An Event Camera Target Tracking Method Based on Improved Kalman Filter
The method integrates occlusion detection and multi-scale Kalman filtering to address target tracking issues in event cameras, enhancing tracking stability and accuracy in complex environments.
Patent Information
- Application Number
- CN202510034616.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The event camera target tracking method based on Kalman filtering is prone to mis-tracking and miss tracking in complex backgrounds, especially when target occlusion and scale changes, it is difficult to achieve stable tracking.
Combining the occlusion detection and recovery mechanism and multi-scale Kalman filtering strategy, the support vector machine model is trained to judge the occlusion through the Gaussian radial basis function, the target trajectory is predicted using a recurrent neural network, and the Kalman filter is dynamically selected according to the target scale for tracking.
It effectively solves the tracking failure problem caused by occlusion and scale changes, realizes stable target tracking in complex backgrounds, improves success rate and accuracy, and has a simple algorithm and small calculation amount.
Smart Images

Figure CN120031920B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of event camera target tracking, and relates to an event camera target tracking method based on improved Kalman filtering. Background Art
[0002] The target tracking methods based on event cameras are mainly divided into two categories: deep learning methods and traditional methods. The deep learning-based event camera target tracking methods achieve accurate tracking of targets in high-dynamic scenes by efficiently extracting features and modeling time series from event streams. These methods usually include using a convolutional neural network (CNN) to extract spatio-temporal features from the event stream, using a recurrent neural network (RNN) or a long short-term memory network (LSTM) to capture the motion trajectory of the target, using a Transformer model with self-attention mechanism to process global features, and using a generative adversarial network (GAN) to improve the robustness of target tracking.
[0003] In traditional methods, the method based on Kalman filtering has been widely used in event target tracking. Kalman filtering is an optimal recursive filter that can estimate the state of a system in the presence of noise and uncertainty. In target tracking applications, Kalman filtering can be used to predict and update state variables such as the position and velocity of the target. Compared with modern machine learning methods such as deep learning, it does not rely on large-scale training data and complex neural network structures.
[0004] The target tracking methods based on event cameras show their respective advantages and limitations in different application scenarios. Deep learning methods perform well in dealing with complex scenarios and multi-target tracking, but due to their high computational complexity, they require a large amount of training data and computing resources, which may become a bottleneck in resource-constrained embedded systems and are not suitable for real-time applications. On the contrary, traditional methods such as Kalman filtering have significant advantages in terms of computational efficiency and real-time performance, and are suitable for real-time applications and resource-constrained embedded systems. However, in complex backgrounds, such as when encountering problems like target occlusion and scale change, it is prone to false tracking and missed tracking phenomena, which affects its reliability and accuracy in complex scenarios. Summary of the Invention
[0005] In order to solve the technical problems of false tracking and missed tracking caused by frequent occlusion and scale transformation in the target tracking of event cameras based on Kalman filtering in complex backgrounds, the present invention proposes an event camera target tracking method based on improved Kalman filtering. This method tightly combines an occlusion detection and recovery mechanism and a multi-scale Kalman filtering strategy, thereby effectively coping with the problems of false tracking and missed tracking caused by frequent occlusion and scale transformation, and achieving continuous and stable tracking when target occlusion and scale change occur in complex backgrounds.
[0006] The object of the present invention is specifically realized through the following technical solutions:
[0007] The present invention discloses an event camera target tracking method based on improved Kalman filtering, and the method includes:
[0008] Step 1, accumulating events in the received event stream according to a time window with a fixed length, and aggregating the event data within each time window into a two-dimensional event frame;
[0009] Step 2, training a support vector machine model through Gaussian radial basis functions, and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combining the feature information before occlusion in the two-dimensional event frame, using a recurrent neural network to predict the target trajectory during occlusion, and after the occlusion ends, restoring the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute Step 3;
[0010] Step 3, creating a Kalman filter corresponding to the target scale and initializing each Kalman filter; predicting each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtaining the Kalman gain based on the covariance prediction of the current state, and updating each predicted Kalman filter through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter;
[0011] Step 4, dynamically selecting the improved Kalman filter with the corresponding scale according to the scale change of the target for event camera target tracking.
[0012] In Step 1, the method of accumulating events in the received event stream according to a time window with a fixed length and aggregating the event data within each time window into a two-dimensional event frame includes:
[0013] Creating an event buffer, accumulating events in the received event stream according to a preset time window with a fixed length, and taking the two-dimensional array formed by all event data at each pixel position within each window as the two-dimensional event frame.
[0014] In Step 2, the method of training a support vector machine model through Gaussian radial basis functions and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes:
[0015] Labeling the occlusion labels and non-occlusion labels for the data set according to the occlusion situation of the target in the data set;
[0016] Using the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame;
[0017] Using the sum of edge features and the labeled occlusion labels, select a Gaussian radial basis function to train a support vector machine model;
[0018] Use the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame.
[0019] In step two, using the Canny edge detection algorithm, the method for calculating the sum of edge features of the target area in the input image or event frame includes:
[0020] 1), Perform Gaussian filtering on the input image or event frame to obtain each pixel point in the Gaussian filter template; among them, the calculation formula for each pixel point in the Gaussian filter template is:
[0021]
[0022] where f(x,y) is each pixel point in the Gaussian filter template, σ is the Gaussian kernel standard deviation, and x and y are the coordinates of a certain position in the Gaussian filter template relative to the center of the Gaussian filter template;
[0023] 2), Use the Sobel operator to return the partial derivatives in the horizontal G x direction and the vertical G y direction to obtain the gradient intensity G and direction θ of each pixel point in the Gaussian filter template:
[0024]
[0025] 3), Traverse all pixel points. If the gradient intensity of a pixel point is greater than or equal to the threshold compared with two pixels along its gradient direction, retain the gradient intensity; if the gradient intensity of a pixel point is less than the threshold compared with two pixels along its gradient direction, set its gradient intensity to 0;
[0026] 4), Set a high threshold and a low threshold. Mark the pixels with gradient intensity higher than the high threshold as edges, and mark the pixels lower than the low threshold as non-edges; if the gradient intensity is both greater than the low threshold and less than the high threshold, and there are edge pixels greater than the high threshold among the adjacent pixels, mark it as an edge; if there are no edge pixels greater than the high threshold among the adjacent pixels, mark it as a non-edge;
[0027] 5), Accumulate the gradient intensities of all edge pixels in the target area to obtain the sum of edge features.
[0028] In step two, using the sum of edge features and the labeled occlusion labels, the method for selecting a Gaussian radial basis function to train a support vector machine model includes:
[0029] Create a support vector machine model using the svm module in the Scikit-learn library in Python. Use the sum of edge features and the labeled occlusion labels, select the Gaussian radial basis function to separate the data points of the occlusion categories to be detected and the non-occlusion categories by finding an optimal hyperplane or hypersurface, and train the support vector machine model; among them, the Gaussian radial basis function is:
[0030]
[0031] Among them, K(m,n) is the value of the Gaussian radial basis function, γ is the kernel parameter, ||m - n|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of two samples to be calculated for the Euclidean distance in the training set respectively.
[0032] In step two, the method of using the trained support vector machine model to judge occlusion for the input two-dimensional event frame includes:
[0033] Input the two-dimensional event frame into the trained support vector machine model for binary classification. If the output is 0, it is judged as the non-occlusion category; if the output is 1, it is judged as the occlusion category.
[0034] In step two, if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, use a recurrent neural network to predict the target trajectory during occlusion, and after occlusion ends, the method of restoring the position of the target according to the predicted target trajectory during occlusion includes:
[0035] a) Label the position of the bounding box of the target in each video frame in the dataset, and divide it into a training set and a test set according to a preset ratio;
[0036] b) Use the ResNet50 to be trained to extract image features, and remove the last fully connected layer to obtain a high-level feature map;
[0037] c) Construct a recurrent neural network. The input of the recurrent neural network is a sequence of feature maps, and each time step corresponds to the feature map of a video frame; the output is the predicted trajectory, including the position and movement speed of the target; input the obtained high-level feature map into the recurrent neural network, process the feature map through time series, and capture the changes and movement patterns of the target over time;
[0038] d) Load the training set, set the loss function as the mean squared error loss, and the optimizer as the adaptive moment estimation optimizer; through multiple rounds of iterative training, optimize the loss function, continuously adjust the parameters of the changes and movement patterns of the target over time to minimize the error between the predicted trajectory and the true trajectory, obtain the trained recurrent neural network and test it through the test set;
[0039] e) During the tracking process, when it is detected that the target is occluded, input the preset number of frames of images before occlusion into the trained recurrent neural network to predict the target trajectory during occlusion;
[0040] f) According to the predicted target trajectory, restore the position of the target after the occlusion ends.
[0041] In step three, create a Kalman filter corresponding to the target scale, and the initialization of each Kalman filter includes:
[0042] Create a Kalman filter corresponding to the threshold range according to whether the number of pixels occupied by the target scale is within the set threshold range;
[0043] Initialize the state matrix, state transition matrix, observation matrix, state covariance matrix, observation noise covariance matrix, and process noise covariance matrix of each filter; among them,
[0044] The state matrix X represents the state of the target, including the position (x1, y1) of the target and the velocity components (v x , v y ), and is expressed as:
[0045] X = [x1, y1, v x , v y T ;
[0046] The state transition matrix Φ is used to describe the law of state change over time and is expressed as:
[0047]
[0048] where Δt is the time interval;
[0049] The observation matrix H is used to map the state matrix to the observation space and is expressed as:
[0050]
[0051] The state covariance matrix P is used to describe the uncertainty of state estimation and is expressed as:
[0052]
[0053] The observation noise covariance matrix R is used to describe the noise in the observation value and is set according to the sensor noise characteristics, and is expressed as:
[0054]
[0055] The process noise covariance matrix Q is used to describe the process noise in the state transition process, reflects the randomness of the target movement, and is set according to the movement characteristics of the target, and is expressed as:
[0056]
[0057] In step three, the calculation method of state prediction is as follows:
[0058] X(k|k - 1) = ΦX(k - 1|k - 1);
[0059] where X(k|k - 1) is the system state at time k, Φ is the state transition matrix, and X(k - 1|k - 1) is the optimal result of the current state;
[0060] The calculation method of covariance prediction is as follows:
[0061] P(k|k - 1) = ΦP(k - 1|k - 1)Φ T + Q;
[0062] where P(k|k - 1) is the state covariance matrix corresponding to X(k|k - 1), P(k - 1|k - 1) is the state covariance matrix corresponding to X(k - 1|k - 1), and Q is the process noise covariance matrix.
[0063] In step three, the calculation method of the Kalman gain includes:
[0064]
[0065] where K is the Kalman gain, which is used to adjust the weight between the predicted value and the observed value of the current position; H is the observation matrix, and R is the observation noise covariance matrix;
[0066] The calculation method of the state update at the next moment includes:
[0067] X(k|k) = X(k|k - 1) + K(Z(k)) - HX(k|k - 1);
[0068] where X(k|k) is the state at the next moment, and Z(k) is the observed value at time k;
[0069] The calculation method of the covariance update of the current state includes:
[0070] P(k|k) = (I - KH)P(k|k - 1);
[0071] where P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
[0072] The beneficial effects of the present invention are as follows:
[0073] 1. The present invention converts the event stream into a two - dimensional event frame, which is convenient for subsequent algorithm use and visualization.
[0074] 2. The present invention solves the problem of target loss caused by frequent occlusions in event-based target tracking by setting up an occlusion detection and recovery mechanism. Through sample annotation of the dataset, feature extraction using the Canny edge detection algorithm, and training and application of the support vector machine model, real-time judgment of whether the target is occluded is achieved, ensuring immediate recognition when the target is occluded and avoiding tracking interruption caused by occlusion. During the occlusion recovery process, combined with the feature information before occlusion in the two-dimensional event frame, a recurrent neural network is used to predict the trajectory during occlusion, and the target position is restored based on the predicted trajectory after the occlusion ends, ensuring the continuity of tracking.
[0075] 3. The present invention creates a Kalman filter corresponding to the target scale and initializes each Kalman filter; based on the initialization parameters and the current position of the target, each Kalman filter is predicted to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; based on the covariance prediction of the current state, the Kalman gain is obtained, and each predicted Kalman filter is updated through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter, effectively solving the problem of tracking failure caused by scale changes during the target tracking process, and being able to dynamically select a filter with an appropriate scale for tracking according to the scale change of the target, ensuring stable tracking of target changes at different scales.
[0076] 4. When tested on the event UAV dataset, the average success rate reaches 0.758, and the average precision reaches 0.613, which is higher than the traditional Kalman filter algorithm.
[0077] 5. The algorithm is simple. Compared with the target tracking algorithm based on deep learning, its computational complexity is smaller and the speed is faster. Compared with the traditional target tracking algorithm based on Kalman filter, it has a higher success rate and precision in dealing with occlusion and scale change problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0079] Figure 1 It is a schematic diagram of a two-dimensional event frame in an embodiment of the present invention.
[0080] Figure 2 It is one of the results of target tracking using the improved multi-scale Kalman filter in an embodiment of the present invention.
[0081] Figure 3 It is the second result of target tracking using the improved multi-scale Kalman filter in an embodiment of the present invention.
[0082] Figure 4This is the third result of target tracking using the improved multi-scale Kalman filter in the embodiments of the present invention.
[0083] Figure 5 This is the fourth result of target tracking using the improved multi-scale Kalman filter in the embodiments of the present invention.
[0084] Figure 6 This is the fifth result of target tracking using the improved multi-scale Kalman filter in the embodiments of the present invention.
[0085] Figure 7 This is the sixth result of target tracking using the improved multi-scale Kalman filter in the embodiments of the present invention. Detailed implementation manners
[0086] The embodiments of the present invention provide an event camera target tracking method based on an improved Kalman filter, and the method includes:
[0087] Step 1: Cumulate the received event stream according to a time window with a fixed length, and the event data within each time window is aggregated into a two-dimensional event frame;
[0088] Step 2: Train a support vector machine model through Gaussian radial basis functions, and input the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, and use a recurrent neural network to predict the target trajectory during occlusion. After the occlusion ends, restore the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute Step 3;
[0089] Step 3: Create a Kalman filter corresponding to the target scale and initialize each Kalman filter; based on the initialization parameters and the current position of the target, predict each Kalman filter to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtain the Kalman gain based on the covariance prediction of the current state, and update each predicted Kalman filter through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter;
[0090] Step 4: Dynamically select the improved Kalman filter corresponding to the scale according to the scale change of the target for event camera target tracking.
[0091] In Step 1, the method of cumulating the received event stream according to a time window with a fixed length and aggregating the event data within each time window into a two-dimensional event frame includes:
[0092] Create an event buffer, accumulate the received event stream according to a preset fixed-length time window, and form a two-dimensional array of all event data at each pixel position within each window as a two-dimensional event frame.
[0093] For example: Create an event buffer, divide the event data by a 20ms event window, count the number of events at each pixel position within each window, form a two-dimensional array as the event frame, and visualize the event frame as Figure 1 shown.
[0094] In step two, the method of training a support vector machine model through a Gaussian radial basis function and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes:
[0095] Annotate the occlusion labels and non-occlusion labels for the dataset according to the occlusion situation of the targets in the dataset;
[0096] For example: Use a 6.8k-frame event drone dataset to annotate the occlusion labels and non-occlusion labels, with a total of 2.7k occlusion frames and 4.1k non-occlusion frames annotated;
[0097] Use the Canny edge detection algorithm to calculate the total edge features of the target area in the input image or event frame;
[0098] Use the total edge features and the annotated occlusion labels to select a Gaussian radial basis function to train the support vector machine model;
[0099] Use the trained support vector machine model to judge the occlusion of the input two-dimensional event frame.
[0100] In step two, the method of using the Canny edge detection algorithm to calculate the total edge features of the target area in the input image or event frame includes:
[0101] 1) Perform Gaussian filtering on the input image or event frame to obtain each pixel point in the Gaussian filter template to smooth the image and reduce noise; among them, the calculation formula for each pixel point in the Gaussian filter template is:
[0102]
[0103] Among them, f(x,y) is each pixel point in the Gaussian filter template, σ is the Gaussian kernel standard deviation, and x and y are the coordinates of a certain position in the Gaussian filter template relative to the center of the Gaussian filter template;
[0104] 2) Use the Sobel operator to return the partial derivatives in the horizontal G x direction and the vertical G y direction to obtain the gradient intensity G and direction θ of each pixel point in the Gaussian filter template:
[0105]
[0106] G x = S x * A;
[0107] G y = S y * A;
[0108]
[0109] where A is a time window, set as a 3x3 window in the image, S x is the Sobel operator in the x direction, S y is the Sobel operator in the y direction, and * is the convolution symbol.
[0110] 3), Traverse all pixel points. If the gradient intensity of a pixel point is greater than or equal to the threshold compared with two pixels along its gradient direction, then retain the gradient intensity; if the gradient intensity of a pixel point is less than the threshold compared with two pixels along its gradient direction, then set its gradient intensity to 0; perform non-maximum suppression in the gradient direction to retain local maxima and suppress non-maxima, thereby refining the edges.
[0111] 4), Set a high threshold and a low threshold. Mark the pixels with gradient intensity higher than the high threshold as edges, and mark the pixels lower than the low threshold as non-edges; if the gradient intensity is both greater than the low threshold and less than the high threshold, and there are edge pixels greater than the high threshold among the adjacent pixels, then mark it as an edge; if there are no edge pixels greater than the high threshold among the adjacent pixels, then mark it as a non-edge;
[0112] For example: Set the high threshold to 60 and the low threshold to 20. Mark the pixels with gradient intensity higher than 60 as edges, and mark the pixels lower than 20 as non-edges. If the detection result is both greater than the low threshold and less than the high threshold, then it is necessary to determine whether there are edge pixels greater than the high threshold among the adjacent pixels of the pixel. If so, it is an edge, otherwise it is not.
[0113] 5), Accumulate the gradient intensities of all edge pixels within the target area to obtain the total edge feature.
[0114] In step two, the method of using the total edge feature and the labeled occlusion label to select the Gaussian radial basis function to train the support vector machine model includes:
[0115] Create a support vector machine model using the svm module of the Scikit-learn library in Python. Use the sum of edge features and the labeled occlusion labels. Select the Gaussian radial basis function to separate the data points of the occlusion category to be detected and the non-occlusion category by finding an optimal hyperplane or hypersurface, and train the support vector machine model; among them, the Gaussian radial basis function is:
[0116]
[0117] Among them, K(x,y) is the value of the Gaussian radial basis function, γ is the kernel parameter, for example, set to 1, ||m - n|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of two samples to be calculated for the Euclidean distance in the training set respectively.
[0118] In step two, the method of using the trained support vector machine model to judge occlusion for the input two-dimensional event frame includes:
[0119] Input the two-dimensional event frame into the trained support vector machine model for binary classification. If the output is 0, it is judged as the non-occlusion category; if the output is 1, it is judged as the occlusion category.
[0120] In step two, if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, and use a recurrent neural network to predict the target trajectory during occlusion. After occlusion ends, the method of restoring the position of the target according to the predicted target trajectory during occlusion includes:
[0121] a) Mark the bounding box position of the target in each video frame in the dataset, and divide it into a training set and a test set according to a preset ratio such as 7:3;
[0122] b) Use the ResNet50 to be trained to extract image features, and remove the last fully connected layer to obtain high-level feature maps; these feature maps capture the key features of the target, such as shape and texture, and provide a basis for subsequent sequence processing.
[0123] c) Build a recurrent neural network. The input of the recurrent neural network is a sequence of feature maps, and each time step corresponds to the feature map of a video frame; the output is the predicted trajectory, including the position and movement speed of the target; input the obtained high-level feature maps into the recurrent neural network, and process the feature maps through time series to capture the changes and movement patterns of the target over time;
[0124] d) Load the training set, set the loss function to mean squared error loss, and the optimizer to adaptive moment estimation optimizer; through multiple rounds of iterative training, optimize the loss function, and continuously adjust the changes of the target over time and the motion pattern parameters to minimize the error between the predicted trajectory and the true trajectory, obtaining a trained recurrent neural network and testing it with the test set;
[0125] e) During the tracking process, when it is detected that the target is occluded, input the preset number of frames of images before occlusion into the trained recurrent neural network to predict the target trajectory during occlusion;
[0126] f) According to the predicted target trajectory, restore the position of the target after the occlusion ends. If no occlusion is detected, continue to track the target normally.
[0127] In step three, create a Kalman filter corresponding to the target scale, and initialize each Kalman filter including:
[0128] Create a Kalman filter corresponding to the threshold range according to whether the number of pixels occupied by the target scale is within the set threshold range;
[0129] For example: It is preferably to create three Kalman filters, which are respectively used to process targets of small scale, medium scale, and large scale. When the number of pixels occupied by the target is less than or equal to 10, the target is considered a small-scale target; when the number of pixels occupied by the target is greater than 10 and less than or equal to 60, the target is considered a medium-scale target; when the number of pixels occupied by the target is greater than or equal to 60, it is considered a large-scale target.
[0130] Initialize the state matrix, state transition matrix, observation matrix, state covariance matrix, observation noise covariance matrix, and process noise covariance matrix of each filter; among them,
[0131] The state matrix X represents the state of the target, including the position (x1, y1) of the target and the velocity components (v x , v y ), and is expressed as:
[0132] X = [x1, y1, v x , v y T ;
[0133] The state transition matrix Φ is used to describe the law of change of the state over time and is expressed as:
[0134]
[0135] where Δt is the time interval;
[0136] The observation matrix H is used to map the state matrix to the observation space and is expressed as:
[0137]
[0138] The state covariance matrix P, which is used to describe the uncertainty of state estimation, is expressed as:
[0139]
[0140] The observation noise covariance matrix R, which is used to describe the noise in the observed values and is set according to the sensor noise characteristics, is expressed as:
[0141]
[0142] The process noise covariance matrix Q, which is used to describe the process noise in the state transition process, reflects the randomness of the target motion, and is set according to the motion characteristics of the target, is expressed as:
[0143]
[0144] In step three, the calculation method of state prediction is:
[0145] X(k|k - 1) = ΦX(k - 1|k - 1);
[0146] where X(k|k - 1) is the system state at time k, Φ is the state transition matrix, and X(k - 1|k - 1) is the optimal result of the current state;
[0147] The calculation method of covariance prediction is:
[0148] P(k|k - 1) = ΦP(k - 1|k - 1)Φ T +Q;
[0149] where P(k|k - 1) is the state covariance matrix corresponding to X(k|k - 1), P(k - 1|k - 1) is the state covariance matrix corresponding to X(k - 1|k - 1), and Q is the process noise covariance matrix.
[0150] In step three, the calculation method of the Kalman gain includes:
[0151]
[0152] where K is the Kalman gain, which is used to adjust the weight between the predicted value and the current position observed value; H is the observation matrix, and R is the observation noise covariance matrix;
[0153] The calculation method of the state update at the next moment includes:
[0154] X(k|k) = X(k|k - 1) + K(Z(k)) - HX(k|k - 1);
[0155] Among them, X(k|k) is the state at the next moment, and Z(k) is the observed value at the k-th moment;
[0156] The calculation method for updating the covariance of the current state includes:
[0157] P(k|k) = (I - KH)P(k|k - 1);
[0158] Among them, P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
[0159] Each Kalman filter independently performs prediction and update, and can dynamically select a suitable improved Kalman filter for tracking according to the scale change of the target. In this way, the improved Kalman filter can maintain stable tracking performance at different scales, effectively solving the problem of tracking failure caused by scale change.
[0160] The tracking results are as Figures 2 to 7 shown, which details the target tracking in different scenarios.
[0161] Figure 2 , Figure 3 and Figure 4 show some data of the target in the road scene of the UAV event dataset under occlusion. Figure 2 In, the UAV target flies over the road. In Figure 3 the UAV target is partially occluded by a moving vehicle. Despite the occlusion, through the occlusion detection and recovery mechanism, Figure 3 and Figure 4 the UAV targets in can still achieve continuous and stable tracking, and this result shows the effectiveness and reliability of this mechanism in dealing with occlusion problems.
[0162] Figure 5 , Figure 6 and Figure 7 show the situation in a more complex woods background scene of the UAV event dataset. In this scene, the UAV target faces challenges such as frequent occlusion and scale change. To address these issues, the occlusion detection and recovery mechanism disclosed in the present invention is adopted, and Kalman filters adapted to different scales are dynamically selected and switched during the tracking process. The experimental results show that these methods effectively overcome the problems brought by occlusion and scale change, ensuring continuous and stable tracking of the UAV target.
[0163] The beneficial effects of the embodiments of the present invention are:
[0164] 1. The present invention converts the event stream into a two-dimensional event frame, facilitating subsequent algorithm use and visualization.
[0165] 2. By setting up an occlusion detection and recovery mechanism, the present invention solves the problem of target loss caused by frequent occlusions in event-based target tracking. Through sample annotation of the dataset, feature extraction using the Canny edge detection algorithm, and training and application of the support vector machine model, real-time judgment of whether the target is occluded is achieved, ensuring immediate recognition when the target is occluded and avoiding tracking interruption caused by occlusion. During the occlusion recovery process, combined with the feature information before occlusion in the two-dimensional event frame, a recurrent neural network is used to predict the trajectory during occlusion, and the target position is restored based on the predicted trajectory after the occlusion ends, ensuring the continuity of tracking.
[0166] 3. The present invention creates a Kalman filter corresponding to the target scale and initializes each Kalman filter; based on the initialization parameters and the current position of the target, each Kalman filter is predicted to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; based on the covariance prediction of the current state, the Kalman gain is obtained, and each predicted Kalman filter is updated through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter, which effectively solves the problem of tracking failure caused by scale changes during the target tracking process. It can dynamically select a filter with an appropriate scale for tracking according to the scale change of the target, ensuring stable tracking of target changes at different scales.
[0167] 4. When tested on the event drone dataset, the average success rate reaches 0.758 and the average precision reaches 0.613, which is higher than the traditional Kalman filtering algorithm.
[0168] 5. The algorithm is simple. Compared with the deep learning-based target tracking algorithm, its computational complexity is smaller and the speed is faster. Compared with the traditional Kalman filter-based target tracking algorithm, it has a higher success rate and precision in dealing with occlusion and scale change problems.
[0169] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An event camera target tracking method based on improved Kalman filtering, characterized in that The method includes: Step 1: Accumulate the received event stream according to time windows of a fixed length, and the event data within each time window is aggregated into a two-dimensional event frame; Step 2: Train a support vector machine model through Gaussian radial basis functions, and input the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, and use a recurrent neural network to predict the target trajectory during occlusion. After the occlusion ends, restore the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute Step 3; Step 3: Create a Kalman filter corresponding to the target scale and initialize each Kalman filter; based on the initialization parameters and the current position of the target, predict each Kalman filter to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtain the Kalman gain based on the covariance prediction of the current state, and update each predicted Kalman filter through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter; Step 4: Dynamically select the improved Kalman filter corresponding to the scale according to the scale change of the target for event camera target tracking.
2. The method according to claim 1, wherein In Step 1, the method of accumulating the received event stream according to time windows of a fixed length, and aggregating the event data within each time window into a two-dimensional event frame includes: Create an event buffer, accumulate the received event stream according to time windows of a preset fixed length, and use the two-dimensional array formed by all event data at each pixel position within each window as the two-dimensional event frame.
3. The method according to claim 1 or 2, characterized in that, In Step 2, the method of training a support vector machine model through Gaussian radial basis functions and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes: Label the dataset with occlusion labels and non-occlusion labels according to the occlusion situation of the target in the dataset; Use the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame; Select a Gaussian radial basis function to train a support vector machine model using the sum of edge features and the labeled occlusion labels; Use the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame.
4. The method according to claim 3, characterized in that, In Step 2, the method of using the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame includes: 1) Perform Gaussian filtering on the input image or event frame to obtain each pixel point in the Gaussian filter template; among them, the calculation formula for each pixel point in the Gaussian filter template is: where f(x,y) is each pixel point in the Gaussian filter template, σ is the Gaussian kernel standard deviation, and x and y are the coordinates of a certain position in the Gaussian filter template relative to the center of the Gaussian filter template; 2), use the Sobel operator to return the horizontal G x direction and the vertical G y direction partial derivatives, and obtain the gradient intensity G and direction θ of each pixel point in the Gaussian filter template: 3) Traverse all pixel points. If the gradient intensity of a pixel point is greater than or equal to the threshold compared with two pixels along its gradient direction, retain the gradient intensity; if the gradient intensity of a pixel point is less than the threshold compared with two pixels along its gradient direction, set its gradient intensity to 0; 4), Set a high threshold and a low threshold, mark the pixels with gradient intensity higher than the high threshold as edges, and mark the pixels lower than the low threshold as non-edges; if the gradient intensity is both greater than the low threshold and less than the high threshold, and there are edge pixels greater than the high threshold among the adjacent pixels, then mark it as an edge; if there are no edge pixels greater than the high threshold among the adjacent pixels, then mark it as a non-edge; 5), Accumulate the gradient intensities of all edge pixels within the target area to obtain the total edge feature.
5. The method according to claim 3, characterized in that, In step two, the method of using the total edge feature and the labeled occlusion label to select the Gaussian radial basis function to train the support vector machine model includes: Create a support vector machine model using the svm module of the Scikit-learn library in python. Use the total edge feature and the labeled occlusion label to select the Gaussian radial basis function to separate the data points of the occlusion category to be detected and the non-occlusion category by finding an optimal hyperplane or hypersurface, and train the support vector machine model; among them, the Gaussian radial basis function is: Among them, K(m,n) is the value of the Gaussian radial basis function, γ is the kernel parameter, ||m - n|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of two samples to be calculated for the Euclidean distance in the training set respectively.
6. The method according to claim 3, wherein In step two, the method of using the trained support vector machine model to judge occlusion for the input two-dimensional event frame includes: Input the two-dimensional event frame into the trained support vector machine model for binary classification. If the output is 0, it is judged as the non-occlusion category; if the output is 1, it is judged as the occlusion category.
7. The method according to claim 1, wherein In step two, if occlusion is detected, combined with the feature information before occlusion in the two-dimensional event frame, use a recurrent neural network to predict the target trajectory during occlusion. After the occlusion ends, the method of restoring the position of the target according to the predicted target trajectory during occlusion includes: a) Mark the boundary box position of the target in each video frame of the dataset, and divide it into a training set and a test set according to a preset ratio; b) Use the untrained ResNet50 to extract image features, remove the last fully connected layer to obtain a high-level feature map; c) Construct a recurrent neural network. The input of the recurrent neural network is a sequence of feature maps, and each time step corresponds to the feature map of a video frame; the output is the predicted trajectory, including the position and motion speed of the target; input the obtained high-level feature map into the recurrent neural network, and process the feature map through time series to capture the changes and motion patterns of the target over time; d) Load the training set, set the loss function as the mean squared error loss, and the optimizer as the adaptive moment estimation optimizer; through multiple rounds of iterative training, optimize the loss function, and continuously adjust the parameters of the changes and motion patterns of the target over time to minimize the error between the predicted trajectory and the true trajectory, obtain the trained recurrent neural network and test it through the test set; e) During the tracking process, when it is detected that the target is occluded, input the preset number of frames of images before occlusion into the trained recurrent neural network to predict the target trajectory during occlusion; f) According to the predicted target trajectory, restore the position of the target after the occlusion ends.
8. The method according to claim 1, wherein In step 3, a Kalman filter corresponding to the target scale is created, and the initialization of each Kalman filter includes: Create a Kalman filter corresponding to the threshold range according to whether the pixels occupied by the target scale are within the set threshold range; Initialize the state matrix, state transition matrix, observation matrix, state covariance matrix, observation noise covariance matrix, and process noise covariance matrix of each filter; among them, The state matrix X, representing the state of the target, includes the position (x1, y1) and velocity components (v x , v y ), and is expressed as: X = [x1, y1, v x , v y T ; The state transition matrix Φ, which is used to describe the variation law of the state over time, is expressed as: where Δt is the time interval; The observation matrix H, which is used to map the state matrix to the observation space, is expressed as: The state covariance matrix P, which is used to describe the uncertainty of the state estimation, is expressed as: The observation noise covariance matrix R, which is used to describe the noise in the observed value and is set according to the sensor noise characteristics, is expressed as: The process noise covariance matrix Q, which is used to describe the process noise in the state transition process and reflects the randomness of the target motion, is set according to the motion characteristics of the target, and is expressed as:
9. The method according to claim 1 or 8, characterized in that In step 3, the calculation method of state prediction is: X(k|k - 1) = ΦX(k - 1|k - 1); where X(k|k - 1) is the system state at time k, Φ is the state transition matrix, and X(k - 1|k - 1) is the optimal result of the current state; The calculation method of covariance prediction is: P(k|k - 1) = ΦP(k - 1|k - 1)Φ T + Q; where P(k|k - 1) is the state covariance matrix corresponding to X(k|k - 1), P(k - 1|k - 1) is the state covariance matrix corresponding to X(k - 1|k - 1), and Q is the process noise covariance matrix.
10. The method according to claim 9, characterized in that, In step 3, the calculation method of the Kalman gain includes: where K is the Kalman gain, which is used to adjust the weight between the predicted value and the current position observed value; H is the observation matrix, and R is the observation noise covariance matrix; The calculation method of the state update at the next moment includes: X(k|k) = X(k|k - 1) + K(Z(k)) - HX(k|k - 1); where X(k|k) is the state at the next moment, and Z(k) is the observed value at time k; The calculation method of the covariance update of the current state includes: P(k|k) = (I - KH)P(k|k - 1); where P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
Citation Information
Patent Citations
Anti-shielding tracking method for moving target
CN117036740A
Image tracking method based on filtering algorithm
CN118864509A