Spatial-temporal information perception tracking method based on Kalman filtering
By combining the Kalman filter and the Transformer model, integrating the prediction and update mechanism of the Kalman filter with the feature extraction of the deep learning model, the problem of small targets being lost or misidentified in video sequences is solved, achieving higher tracking accuracy and robustness.
Patent Information
- Application Number
- CN202510918470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-07
Smart Images

Figure CN120912853A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of target tracking, and particularly relates to a spatiotemporal information perception tracking method based on Kalman filtering, which is suitable for small target tracking in computer vision, especially in application scenarios such as video surveillance, robot technology, automatic driving and sports analysis. BACKGROUND
[0002] Target tracking is a key technology in the field of computer vision, which plays a crucial role in multiple application fields such as video surveillance, robot technology, automatic driving and sports analysis. The core task of target tracking is to accurately estimate the position and motion state of target objects in consecutive video frames. With the rapid development of deep learning technology, the field of target tracking has made significant progress, especially in dealing with challenging problems such as occlusion, illumination changes and camera motion. However, despite the achievements of existing research, tracking small targets remains a particularly challenging task. Small targets usually occupy a limited number of pixels in image frames, which makes them prone to loss or misidentification in video sequences, especially when moving quickly. The limited available feature information of small targets makes it difficult for tracking algorithms to maintain stable and accurate tracking. In addition, the high-speed motion of small targets can cause significant displacement between consecutive frames, increasing the difficulty of tracking.
[0003] In real-world application scenarios such as sports analysis, unmanned aerial vehicle monitoring and automatic driving, accurate tracking of small targets is particularly crucial. For example, in sports analysis, accurately tracking athletes and balls can provide key motion analysis data; in unmanned aerial vehicle monitoring, tracking small objects moving quickly is crucial for safety monitoring; and in the field of automatic driving, accurately identifying and tracking small objects on the road is of great significance to avoid accidents and improve driving safety.
[0004] To address these challenges, researchers have explored various methods to improve the tracking performance of small and fast-moving targets. Although Kalman filters show potential in target tracking, how to effectively integrate Kalman filters and fully utilize their capabilities in motion prediction and noise reduction to improve the accuracy and robustness of tracking systems remains an open problem. In addition, how to design deep learning models to extract more robust and discriminative features, and how to combine spatiotemporal information to enhance tracking performance, are also hot topics in current research.
[0005] The present application aims to provide a spatiotemporal information perception tracking method for small targets based on Kalman filtering, which effectively improves the accuracy and robustness of target tracking by combining the prediction and update mechanism of Kalman filters and the powerful feature extraction capability of deep learning models, providing a new solution for the field of visual target tracking. SUMMARY
[0006] The purpose of the present application is to provide a Kalman filter-based spatio-temporal information perception small target tracking method that can effectively solve the problem of insufficient accuracy and robustness in tracking small targets in the prior art.
[0007] The technical solution of the present application is a visual target tracking method combining Kalman filter and Transformer model, which enhances the prediction ability of fast target motion and the robustness to noise by integrating Kalman filter and deep learning technology.
[0008] The Kalman filter-based spatio-temporal information perception tracking process is shown in Figure 1 The framework of KSTrack model is shown in Figure 2 In the initialization phase, the target template is extracted from the first frame of the video sequence and input into the KSTrack model to generate the initial feature representation of the target. At the same time, the state vector and covariance matrix of the Kalman filter are initialized based on the initial bounding box of the target. The state vector includes the position and size of the target, while the covariance matrix represents the uncertainty of the initial state.
[0009] For the processing of each frame, the Kalman filter is used to predict the target state of the current frame based on the state of the previous frame. The prediction step can be represented as:
[0010] (1)
[0011] (2)
[0012] where, is the predicted state at time is the covariance matrix of the predicted state.
[0013] The search region is extracted from the current frame image and input into the KSTrack model together with the template feature to extract the features of the target. The KSTrack model is based on the ViT architecture and can learn the complex feature representation of the target, so as to better cope with the appearance changes of the target and background interference. The KSTrack model output includes the predicted bounding box of the target, the confidence score and the response map, which represents the position probability distribution of the target in the search region.
[0014] The Hanning window function is applied to the response map, and the peak position of the response map is calculated to obtain the final predicted bounding box of the target.
[0015] The final predicted bounding box of the target is input into the Kalman filter as a new measurement to update the state of the target. The update step can be represented as:
[0016]
[0017]
[0018]
[0019] wherein, is the Kalman gain, used to weigh the credibility of the predicted value and the measured value; is the measured value at time . is the updated state covariance matrix.
[0020] The present application proposes a Kalman filter-based spatiotemporal information perception small target tracking method. This method aims to overcome the shortcomings of existing small target tracking techniques in terms of accuracy and robustness, by integrating Kalman filter and deep learning model, significantly improving the tracking ability of small targets in complex environments, providing a new and efficient solution for the small target tracking field. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 Kalman filter-based spatiotemporal information perception tracking process
[0022] Figure 2 Kalman filter-based spatiotemporal information perception tracking model KSTrack architecture
[0023] Figure 3 Examples of comparison results of KSTrack model and other models on small target dataset DETAILED DESCRIPTION
[0024] The present application is a Kalman filter-based spatiotemporal information perception tracking method, the specific steps are as follows:
[0025] (1) First, prepare the public dataset required for training and testing the framework, including TSFMO, Got-10k, LaSOT, COCO, and TrackingNet, etc.
[0026] (2) Construct the spatiotemporal information perception tracking framework KSTrack, and use the network model based on ViT as the backbone network to extract the features of the template image and search image with sizes of 128x128 and 256x256 input into the framework.
[0027] (3) Apply a spatial encoder to the target template and search area to generate spatial feature maps. Input these spatial feature maps into a temporal decoder to capture the temporal dynamic characteristics of the target in consecutive frames.
[0028] (4) The temporal features output by the temporal decoder are fused with the spatial features output by the spatial encoder to form a feature representation that is rich in both temporal and spatial information. This is used to integrate feature information from different time points and spatial locations, enhancing the recognizability of the target.
[0029] (5) The KSTrack model outputs include predicted bounding boxes, confidence scores, and response maps. The response map represents the probability distribution of the target's location in the search area. A Hann window function is applied to the response map to suppress edge effects and improve prediction accuracy.
[0030] (6) The Kalman filter predicts the target state of the current frame based on the state of the previous frame. The prediction step includes state transition and covariance update, where the state transition is controlled by the state transition matrix, which describes how the target state transitions from the previous frame to the current frame.
[0031] (7) The KSTrack model is used to track small targets in video. Figure 3 Nine tracking result examples are given, with straight rectangular boxes indicating the predicted bounding boxes of small targets tracked by the KSTrack model.
Claims
1. A Kalman filter based spatio-temporal information aware tracking method, mainly characterized by, Specifically comprising the following steps: S1: In the first frame of the video sequence, a template of the target is extracted according to the given target bounding box and input into the model KSTrack to generate the initial feature representation of the target, and the state and covariance matrix of the Kalman filter are initialized; S2: The backbone network uses a ViT-based network model, uses a spatial encoder to extract features from the template and search area, and generates a spatial feature map; a temporal decoder decodes the feature map output by the spatial encoder to capture the dynamic changes of the target in the time sequence; then, the temporal decoding features from the template and the search area are fused to enhance the spatio-temporal information of the target; S3: In each frame of image, the current state of the target is predicted based on the state information of the previous frame by using the Kalman filter, so as to obtain the predicted position of the target; S3.1: The prediction step comprises predicting the state at the next time instant from the current state estimate and the state transition matrix; in particular, assuming that the state vector of the target is , the state transition matrix is , and the process noise covariance matrix is , the prediction process can be expressed as: (1) (2) wherein, is the predicted state at time is the predicted state, is the covariance matrix of the predicted state; S4: The model KSTrack outputs the predicted bounding box and confidence score of the target, applies the Hann window function on the response map, and calculates the peak position of the response map to obtain the final predicted bounding box of the target; S5: The final predicted bounding box of the target is input into the Kalman filter as a new measurement to update the state of the target; S5.1: The update step includes revising the prediction results to obtain more accurate state estimates in combination with the new measurement data; specifically, assuming the measurement matrix is , and the measurement noise covariance matrix is , the update process can be expressed as: (3) (4) (5) wherein, is a Kalman gain that weighs the reliability of the predicted value and the measured value; is a measured value at time ; and is an updated state covariance matrix; S6: The overall loss function during model training is composed of a classification loss function , an IoU regression loss function, and an L1 regression loss function, as defined in formula (1), wherein and are both constants: (6) S6.1: , , are defined by equations (7), (8), (9), respectively. (7) (8) (9) wherein, in formula (7) is a target prediction value, is an adjustment factor; in formula (8) represents an area of a graphic intersection of a prediction frame and a true value frame, represents an area of a graphic union of a prediction frame and a true value frame; in formula (9) is a target value, is an estimated value, and n is a sample number.