Event camera target tracking method based on improved Kalman filtering
By improving the Kalman filtering method, combining the occlusion detection and recovery mechanism and the multi-scale Kalman filtering strategy, the error tracking and missed tracking problems caused by occlusion and scale changes in event camera target tracking are solved, and the stable tracking effect is achieved in complex backgrounds.
Patent Information
- Application Number
- CN202510034616.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The event camera target tracking method based on Kalman filtering is prone to mis-tracking and miss tracking in complex backgrounds, especially when the target is frequently occluded and scale changes.
An improved Kalman filtering method is adopted, combining the occlusion detection and recovery mechanism and a multi-scale Kalman filtering strategy. The support vector machine model is trained through Gaussian radial basis function to make occlusion judgments, and the recurrent neural network is used to predict the target trajectory during occlusion. Dynamically select the Kalman filter of the appropriate scale for tracking when the scale changes.
It effectively solves the problem of mistracking and missed tracking caused by occlusion and scale changes, and achieves continuous and stable target tracking in complex contexts. The algorithm is simple, has small calculations, is fast, and has higher success rate and accuracy when dealing with occlusion and scale changes.
Smart Images

Figure CN120031920A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of event camera target tracking, and relates to an event camera target tracking method based on improved Kalman filtering. Background Art
[0002] Event camera-based target tracking methods are mainly divided into two categories: deep learning methods and traditional methods. Event camera target tracking methods based on deep learning achieve accurate tracking of targets in high-dynamic scenes by performing efficient feature extraction and time series modeling on event streams. These methods usually include using convolutional neural networks (CNN) to extract spatiotemporal features in event streams, using recurrent neural networks (RNN) or long short-term memory networks (LSTM) to capture the motion trajectory of targets, using the Transformer model with a self-attention mechanism to process global features, and improving the robustness of target tracking through generative adversarial networks (GAN).
[0003] Among traditional methods, the Kalman filter-based method has been widely used in event target tracking. Kalman filter is an optimal recursive filter that can estimate the state of a system in the presence of noise and uncertainty. In target tracking applications, Kalman filter can be used to predict and update state variables such as the position and velocity of the target. Compared with modern machine learning methods such as deep learning, it does not rely on large-scale training data and complex neural network structures.
[0004] Event camera-based target tracking methods show their own advantages and limitations in different application scenarios. Deep learning methods perform well in dealing with complex scenes and multi-target tracking, but due to their high computational complexity, they require a large amount of training data and computing resources, which may become a bottleneck in resource-constrained embedded systems and are not suitable for real-time applications. On the contrary, traditional methods such as Kalman filtering have significant advantages in computational efficiency and real-time performance, and are suitable for real-time applications and resource-constrained embedded systems. However, in complex backgrounds, when encountering problems such as target occlusion and scale changes, they are prone to mistracking and missed tracking, which affects their reliability and accuracy in complex scenes. Summary of the invention
[0005] In order to solve the technical problems of mistracking and missed tracking caused by frequent occlusion and scale change in event camera target tracking under complex background in Kalman filter-based methods, the present invention proposes an event camera target tracking method based on improved Kalman filtering. The method closely combines occlusion detection and recovery mechanism with multi-scale Kalman filtering strategy, thereby effectively coping with the mistracking and missed tracking problems caused by frequent occlusion and scale change, and can achieve continuous and stable tracking when target occlusion and scale change occur under complex background.
[0006] The purpose of the present invention is specifically achieved through the following technical solutions:
[0007] The present invention discloses an event camera target tracking method based on improved Kalman filtering, the method comprising:
[0008] Step 1: accumulate the received event stream according to a fixed-length time window, and aggregate the event data in each time window into a two-dimensional event frame;
[0009] Step 2: Train the support vector machine model through Gaussian radial basis function, input the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, use recurrent neural network to predict the target trajectory during occlusion, and after the occlusion ends, restore the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute step 3;
[0010] Step 3: Create a Kalman filter corresponding to the target scale and initialize each Kalman filter; predict each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtain the Kalman gain based on the covariance prediction of the current state, and update each Kalman filter after the prediction through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter;
[0011] Step 4: Dynamically select an improved Kalman filter of corresponding scale according to the scale change of the target to perform event camera target tracking.
[0012] In step 1, the received event stream is accumulated according to a time window of a fixed length, and the event data in each time window is aggregated into a two-dimensional event frame, including:
[0013] Create an event buffer, accumulate the received event stream according to a preset fixed-length time window, and use the two-dimensional array formed by all event data at each pixel position in each window as a two-dimensional event frame.
[0014] In step 2, the method of training a support vector machine model by using a Gaussian radial basis function and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes:
[0015] Label the dataset with occlusion labels and non-occlusion labels according to the occlusion conditions of the targets in the dataset;
[0016] Use the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame;
[0017] Using the sum of edge features and the annotated occlusion labels, the Gaussian radial basis function is selected to train the support vector machine model;
[0018] Use the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame.
[0019] In step 2, using the Canny edge detection algorithm, the method for calculating the sum of edge features of the target area in the input image or event frame includes:
[0020] 1) Perform Gaussian filtering on the input image or event frame to obtain each pixel point in the Gaussian filter template; wherein the calculation formula for each pixel point in the Gaussian filter template is:
[0021]
[0022] Where f(x,y) is the pixel point in the Gaussian filter template, σ is the standard deviation of the Gaussian kernel, and x and y are the coordinates of a certain position of the Gaussian filter template relative to the center of the Gaussian filter template;
[0023] 2) Use the Sobel operator to return the level G x Direction and vertical G y The partial derivative of the direction obtains the gradient strength G and direction θ of each pixel in the Gaussian filter template:
[0024]
[0025] 3) Traverse all pixels. If the gradient strength of the pixel is greater than or equal to the threshold compared with the two pixels along the gradient direction, the gradient strength is retained; if the gradient strength of the pixel is less than the threshold compared with the two pixels along the gradient direction, the gradient strength is set to 0;
[0026] 4) Set a high threshold and a low threshold, mark pixels with gradient strength higher than the high threshold as edges, and pixels with gradient strength lower than the low threshold as non-edges; if the gradient strength is both greater than the low threshold and less than the high threshold, and there are edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as an edge; if there are no edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as a non-edge;
[0027] 5) Accumulate the gradient strengths of all edge pixels in the target area to obtain the sum of edge features.
[0028] In step 2, the method of using the edge feature sum and the annotated occlusion label to select the Gaussian radial basis function to train the support vector machine model includes:
[0029] Use the svm module of the Scikit-learn library in Python to create a support vector machine model. Use the sum of edge features and the annotated occlusion labels, select the Gaussian radial basis function to find an optimal hyperplane or hypersurface to separate the data points of the occlusion category to be detected from the non-occlusion category, and train the support vector machine model. The Gaussian radial basis function is:
[0030]
[0031] Where K(m,n) is the Gaussian radial basis function value, γ is the kernel parameter, ||mn|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of the two samples in the training set whose Euclidean distance is to be calculated.
[0032] In step 2, the method of using the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame includes:
[0033] The two-dimensional event frame is input into the trained support vector machine model for binary classification. If the output is 0, it is judged as a non-occlusion category; if the output is 1, it is judged as an occlusion category.
[0034] In step 2, if occlusion is detected, the target trajectory during the occlusion period is predicted using a recurrent neural network in combination with feature information before the occlusion in the two-dimensional event frame. After the occlusion ends, the method for restoring the position of the target according to the predicted target trajectory during the occlusion period includes:
[0035] a) Annotate the bounding box position of the target in each video frame in the dataset and divide it into training set and test set according to the preset ratio;
[0036] b) Use the ResNet50 to be trained to extract image features and remove the last fully connected layer to obtain a high-level feature map;
[0037] c) constructing a recurrent neural network, wherein the input of the recurrent neural network is a feature map sequence, and each time step corresponds to a feature map of a video frame; the output is the output predicted trajectory, including the position and movement speed of the target; the obtained high-level feature map is input into the recurrent neural network, and the feature map is processed by the time series to capture the temporal changes and movement patterns of the target;
[0038] d) Load the training set, set the loss function to mean square error loss, and the optimizer to adaptive moment estimation optimizer; optimize the loss function through multiple rounds of iterative training, continuously adjust the temporal changes of the target and the motion pattern parameters to minimize the error between the predicted trajectory and the actual trajectory, obtain the trained recurrent neural network and test it through the test set;
[0039] e) During the tracking process, when it is detected that the target is occluded, a preset number of frames of images before the occlusion are input into the trained recurrent neural network to predict the target trajectory during the occlusion period;
[0040] f) Based on the predicted target trajectory, restore the target position after the occlusion ends.
[0041] In step 3, a Kalman filter corresponding to the target scale is created, and each Kalman filter is initialized including:
[0042] According to whether the pixels occupied by the target scale are within the set threshold range, a Kalman filter corresponding to the threshold range is created;
[0043] Initialize the state matrix, state transfer matrix, observation matrix, state covariance matrix, observation noise covariance matrix and process noise covariance matrix of each filter; where,
[0044] The state matrix X represents the state of the target, including the position of the target (x 1 ,y 1 ) and velocity component (v x ,v y ), expressed as:
[0045] X=[x 1 ,y 1 ,v x ,v y ] T ;
[0046] The state transfer matrix Φ is used to describe the change of state over time and is expressed as:
[0047]
[0048] Where Δt is the time interval;
[0049] The observation matrix H is used to map the state matrix to the observation space and is expressed as:
[0050]
[0051] The state covariance matrix P, used to describe the uncertainty of state estimation, is expressed as:
[0052]
[0053] The observation noise covariance matrix R, which is used to describe the noise in the observations, is set according to the sensor noise characteristics and is expressed as:
[0054]
[0055] The process noise covariance matrix Q is used to describe the process noise in the state transfer process, reflecting the randomness of the target motion. It is set according to the motion characteristics of the target and is expressed as:
[0056]
[0057] In step 3, the calculation method of state prediction is:
[0058] X(k|k-1)=ΦX(k-1|k-1);
[0059] Among them, X(k|k-1) is the system state at time k, Φ is the state transfer matrix, and X(k-1|k-1) is the optimal result of the current state;
[0060] The covariance forecast is calculated as:
[0061] P(k|k-1)=ΦP(k-1|k-1)Φ T +Q;
[0062] Among them, P(k|k-1) is the state covariance matrix corresponding to X(k|k-1), P(k-1|k-1) is the state covariance matrix corresponding to X(k-1|k-1), and Q is the process noise covariance matrix.
[0063] In step 3, the calculation method of Kalman gain includes:
[0064]
[0065] Among them, K is the Kalman gain, which is used to adjust the weight between the predicted value and the current position observation value; H is the observation matrix, and R is the observation noise covariance matrix;
[0066] The calculation method of the state update at the next moment includes:
[0067] X(k|k)=X(k|k-1)+K(Z(k))-HX(k|k-1);
[0068] Among them, X(k|k) is the state at the next moment, and Z(k) is the observation value at moment k;
[0069] The calculation method of the covariance update of the current state includes:
[0070] P(k|k)=(I-KH)P(k|k-1);
[0071] Among them, P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
[0072] The beneficial effects of the present invention are:
[0073] 1. The present invention converts the event stream into a two-dimensional event frame, which is convenient for subsequent algorithm use and visualization.
[0074] 2. The present invention solves the problem of target loss caused by frequent occlusion in event-based target tracking by setting up an occlusion detection and recovery mechanism. Through sample annotation of the data set, feature extraction of the Canny edge detection algorithm, and training and application of the support vector machine model, real-time judgment of whether the target is occluded is achieved, ensuring that the target can be immediately identified when it is occluded, avoiding tracking interruptions caused by occlusion. In the occlusion recovery process, the feature information before occlusion in the two-dimensional event frame is combined, and the recurrent neural network is used to predict the trajectory during occlusion, and the target position is restored according to the predicted trajectory after the occlusion ends to ensure the continuity of tracking.
[0075] 3. The present invention creates a Kalman filter corresponding to the target scale and initializes each Kalman filter; predicts each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtains the Kalman gain based on the covariance prediction of the current state, and updates each Kalman filter after the prediction through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter, which effectively solves the tracking failure problem caused by scale changes in the target tracking process, and can dynamically select a filter of a suitable scale for tracking according to the scale change of the target, ensuring that target changes at different scales can be stably tracked.
[0076] 4. Tested on the event drone dataset, the average success rate reached 0.758 and the average accuracy reached 0.613, which are higher than the traditional Kalman filter algorithm.
[0077] 5. The algorithm is simple, and compared with the target tracking algorithm based on deep learning, it has less computation and faster speed. Compared with the traditional target tracking algorithm based on Kalman filtering, it has a higher success rate and accuracy in dealing with occlusion and scale change problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] The present invention is further described in detail below based on the accompanying drawings and embodiments.
[0079] Figure 1 It is a schematic diagram of a two-dimensional event frame in an example of the present invention.
[0080] Figure 2 This is the target tracking result one of the improved multi-scale Kalman filter in the example of the present invention.
[0081] Figure 3 This is the second target tracking result using the improved multi-scale Kalman filter in the example of the present invention.
[0082] Figure 4 This is the third result of target tracking using the improved multi-scale Kalman filter in the example of the present invention.
[0083] Figure 5 This is the fourth target tracking result of the improved multi-scale Kalman filter in the example of the present invention.
[0084] Figure 6 This is the fifth target tracking result of the improved multi-scale Kalman filter in the example of the present invention.
[0085] Figure 7 This is the sixth result of target tracking using the improved multi-scale Kalman filter in the example of the present invention. DETAILED DESCRIPTION
[0086] An embodiment of the present invention provides an event camera target tracking method based on improved Kalman filtering, the method comprising:
[0087] Step 1: accumulate the received event stream according to a fixed-length time window, and aggregate the event data in each time window into a two-dimensional event frame;
[0088] Step 2: Train the support vector machine model through Gaussian radial basis function, input the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, use recurrent neural network to predict the target trajectory during occlusion, and after the occlusion ends, restore the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute step 3;
[0089] Step 3: Create a Kalman filter corresponding to the target scale and initialize each Kalman filter; predict each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtain the Kalman gain based on the covariance prediction of the current state, and update each Kalman filter after the prediction through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter;
[0090] Step 4: Dynamically select an improved Kalman filter of corresponding scale according to the scale change of the target to perform event camera target tracking.
[0091] In step 1, the received event stream is accumulated according to a time window of a fixed length, and the event data in each time window is aggregated into a two-dimensional event frame, including:
[0092] Create an event buffer, accumulate the received event stream according to a preset fixed-length time window, and use the two-dimensional array formed by all event data at each pixel position in each window as a two-dimensional event frame.
[0093] For example, create an event buffer, divide the event data into 20ms event windows, count the number of events at each pixel position in each window, form a two-dimensional array as an event frame, and visualize the event frame as follows: Figure 1 shown.
[0094] In step 2, the method of training a support vector machine model by using a Gaussian radial basis function and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes:
[0095] Label the dataset with occlusion labels and non-occlusion labels according to the occlusion conditions of the targets in the dataset;
[0096] For example, using a 6.8k-frame event drone dataset to annotate occlusion labels and non-occlusion labels, a total of 2.7k occlusion frames and 4.1k non-occlusion frames are annotated;
[0097] Use the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame;
[0098] Using the sum of edge features and the annotated occlusion labels, the Gaussian radial basis function is selected to train the support vector machine model;
[0099] Use the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame.
[0100] In step 2, using the Canny edge detection algorithm, the method for calculating the sum of edge features of the target area in the input image or event frame includes:
[0101] 1) Perform Gaussian filtering on the input image or event frame to obtain each pixel in the Gaussian filter template to smooth the image and reduce noise; wherein, the calculation formula for each pixel in the Gaussian filter template is:
[0102]
[0103] Where f(x,y) is the pixel point in the Gaussian filter template, σ is the standard deviation of the Gaussian kernel, and x and y are the coordinates of a certain position of the Gaussian filter template relative to the center of the Gaussian filter template;
[0104] 2) Use the Sobel operator to return the level G x Direction and vertical G y The partial derivative of the direction obtains the gradient strength G and direction θ of each pixel in the Gaussian filter template:
[0105]
[0106] G x =S x *A;
[0107] G y =S y *A;
[0108]
[0109] Among them, A is the time window, which is set to a 3x3 window in the image, and S x is the Sobel operator in the x direction, S y is the Sobel operator in the y direction, and * is the convolution symbol.
[0110] 3) Traverse all pixels. If the gradient strength of the pixel is greater than or equal to the threshold compared with the two pixels along the gradient direction, the gradient strength is retained; if the gradient strength of the pixel is less than the threshold compared with the two pixels along the gradient direction, the gradient strength is set to 0; perform non-maximum suppression in the gradient direction to retain the local maximum and suppress the non-maximum, thereby refining the edge.
[0111] 4) Set a high threshold and a low threshold, mark pixels with gradient strength higher than the high threshold as edges, and pixels with gradient strength lower than the low threshold as non-edges; if the gradient strength is both greater than the low threshold and less than the high threshold, and there are edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as an edge; if there are no edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as a non-edge;
[0112] For example, set the high threshold to 60 and the low threshold to 20. If the pixel with gradient intensity higher than 60 is marked as edge, and the pixel with gradient intensity lower than 20 is marked as non-edge. If the detection result is greater than the low threshold and less than the high threshold, it is necessary to determine whether there is an edge pixel greater than the high threshold among the adjacent pixels of the pixel. If so, it is an edge, otherwise it is not.
[0113] 5) Accumulate the gradient strengths of all edge pixels in the target area to obtain the sum of edge features.
[0114] In step 2, the method of using the edge feature sum and the annotated occlusion label to select the Gaussian radial basis function to train the support vector machine model includes:
[0115] Use the svm module of the Scikit-learn library in Python to create a support vector machine model. Use the sum of edge features and the annotated occlusion labels, select the Gaussian radial basis function to find an optimal hyperplane or hypersurface to separate the data points of the occlusion category to be detected from the non-occlusion category, and train the support vector machine model. The Gaussian radial basis function is:
[0116]
[0117] Wherein, K(x, y) is the Gaussian radial basis function value, γ is the kernel parameter, for example, it is set to 1, ||mn|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of the two samples in the training set whose Euclidean distance is to be calculated.
[0118] In step 2, the method of using the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame includes:
[0119] The two-dimensional event frame is input into the trained support vector machine model for binary classification. If the output is 0, it is judged as a non-occlusion category; if the output is 1, it is judged as an occlusion category.
[0120] In step 2, if occlusion is detected, the target trajectory during the occlusion period is predicted using a recurrent neural network in combination with feature information before the occlusion in the two-dimensional event frame. After the occlusion ends, the method for restoring the position of the target according to the predicted target trajectory during the occlusion period includes:
[0121] a) Annotate the bounding box position of the target in each video frame in the dataset and divide it into a training set and a test set according to a preset ratio, such as 7:3;
[0122] b) Use the ResNet50 to be trained to extract image features and remove the last fully connected layer to obtain high-level feature maps; these feature maps capture the key features of the target, such as shape and texture, and provide a basis for subsequent sequence processing.
[0123] c) constructing a recurrent neural network, wherein the input of the recurrent neural network is a feature map sequence, and each time step corresponds to a feature map of a video frame; the output is the output predicted trajectory, including the position and movement speed of the target; the obtained high-level feature map is input into the recurrent neural network, and the feature map is processed by the time series to capture the temporal changes and movement patterns of the target;
[0124] d) Load the training set, set the loss function to mean square error loss, and the optimizer to adaptive moment estimation optimizer; optimize the loss function through multiple rounds of iterative training, continuously adjust the temporal changes of the target and the motion pattern parameters to minimize the error between the predicted trajectory and the actual trajectory, obtain the trained recurrent neural network and test it through the test set;
[0125] e) During the tracking process, when it is detected that the target is occluded, a preset number of frames of images before the occlusion are input into the trained recurrent neural network to predict the target trajectory during the occlusion period;
[0126] f) Based on the predicted target trajectory, restore the target’s position after the occlusion ends. If no occlusion is detected, continue tracking the target normally.
[0127] In step 3, a Kalman filter corresponding to the target scale is created, and each Kalman filter is initialized including:
[0128] According to whether the pixels occupied by the target scale are within the set threshold range, a Kalman filter corresponding to the threshold range is created;
[0129] For example, it is preferred to create three Kalman filters, which are used to process small-scale, medium-scale and large-scale targets respectively. When the number of pixels occupied by the target is less than or equal to 10, the target is considered to be a small-scale target; when the number of pixels occupied by the target is greater than 10 and less than or equal to 60, the target is considered to be a medium-scale target; when the number of pixels occupied by the target is greater than or equal to 60, it is considered to be a large-scale target.
[0130] Initialize the state matrix, state transfer matrix, observation matrix, state covariance matrix, observation noise covariance matrix and process noise covariance matrix of each filter; where,
[0131] The state matrix X represents the state of the target, including the position of the target (x 1 ,y 1 ) and velocity component (v x ,v y ), expressed as:
[0132] X=[x 1 ,y 1 ,v x ,v y ] T ;
[0133] The state transfer matrix Φ is used to describe the change of state over time and is expressed as:
[0134]
[0135] Where Δt is the time interval;
[0136] The observation matrix H is used to map the state matrix to the observation space and is expressed as:
[0137]
[0138] The state covariance matrix P, used to describe the uncertainty of state estimation, is expressed as:
[0139]
[0140] The observation noise covariance matrix R, which is used to describe the noise in the observations, is set according to the sensor noise characteristics and is expressed as:
[0141]
[0142] The process noise covariance matrix Q is used to describe the process noise in the state transfer process, reflecting the randomness of the target motion. It is set according to the motion characteristics of the target and is expressed as:
[0143]
[0144] In step 3, the calculation method of state prediction is:
[0145] X(k|k-1)=ΦX(k-1|k-1);
[0146] Among them, X(k|k-1) is the system state at time k, Φ is the state transfer matrix, and X(k-1|k-1) is the optimal result of the current state;
[0147] The covariance forecast is calculated as:
[0148] P(k|k-1)=ΦP(k-1|k-1)Φ T +Q;
[0149] Among them, P(k|k-1) is the state covariance matrix corresponding to X(k|k-1), P(k-1|k-1) is the state covariance matrix corresponding to X(k-1|k-1), and Q is the process noise covariance matrix.
[0150] In step 3, the calculation method of Kalman gain includes:
[0151]
[0152] Among them, K is the Kalman gain, which is used to adjust the weight between the predicted value and the current position observation value; H is the observation matrix, and R is the observation noise covariance matrix;
[0153] The calculation method of the state update at the next moment includes:
[0154] X(k|k)=X(k|k-1)+K(Z(k))-HX(k|k-1);
[0155] Among them, X(k|k) is the state at the next moment, and Z(k) is the observation value at moment k;
[0156] The calculation method of the covariance update of the current state includes:
[0157] P(k|k)=(I-KH)P(k|k-1);
[0158] Among them, P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
[0159] Each Kalman filter predicts and updates independently, and can dynamically select the appropriate improved Kalman filter for tracking according to the scale change of the target. In this way, the improved Kalman filter can maintain stable tracking performance at different scales, effectively solving the problem of tracking failure caused by scale changes.
[0160] Tracking results such as Figures 2 to 7 As shown in Figure 2, the target tracking in different scenarios is shown in detail.
[0161] Figure 2 , Figure 3 and Figure 4 It shows some data of targets under occlusion in road scenes in the drone event dataset. Figure 2 In the video, a drone target is flying over the road. Figure 3 In the example, the drone target is partially occluded by the moving vehicle. Despite the occlusion, the occlusion detection and recovery mechanism Figure 3 and Figure 4 The UAV target can still be tracked continuously and stably, which shows the effectiveness and reliability of the mechanism in dealing with occlusion problems.
[0162] Figure 5 , Figure 6 and Figure 7 The situation in a more complex forest background scene in the drone event dataset is shown. In this scene, the drone target faces challenges such as frequent occlusion and scale changes. In order to deal with these problems, the occlusion detection and recovery mechanism disclosed in the present invention is adopted, and Kalman filters adapted to different scales are dynamically selected and switched during the tracking process. The experimental results show that these methods effectively overcome the problems caused by occlusion and scale changes, and ensure the continuous and stable tracking of the drone target.
[0163] The beneficial effects of the embodiments of the present invention are:
[0164] 1. The present invention converts the event stream into a two-dimensional event frame, which is convenient for subsequent algorithm use and visualization.
[0165] 2. The present invention solves the problem of target loss caused by frequent occlusion in event-based target tracking by setting up an occlusion detection and recovery mechanism. Through sample annotation of the data set, feature extraction of the Canny edge detection algorithm, and training and application of the support vector machine model, real-time judgment of whether the target is occluded is achieved, ensuring that the target can be immediately identified when it is occluded, avoiding tracking interruptions caused by occlusion. In the occlusion recovery process, the feature information before occlusion in the two-dimensional event frame is combined, and the recurrent neural network is used to predict the trajectory during occlusion, and the target position is restored according to the predicted trajectory after the occlusion ends to ensure the continuity of tracking.
[0166] 3. The present invention creates a Kalman filter corresponding to the target scale and initializes each Kalman filter; predicts each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtains the Kalman gain based on the covariance prediction of the current state, and updates each Kalman filter after the prediction through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter, which effectively solves the tracking failure problem caused by scale changes in the target tracking process, and can dynamically select a filter of a suitable scale for tracking according to the scale change of the target, ensuring that target changes at different scales can be stably tracked.
[0167] 4. Tested on the event drone dataset, the average success rate reached 0.758 and the average accuracy reached 0.613, which are higher than the traditional Kalman filter algorithm.
[0168] 5. The algorithm is simple, and compared with the target tracking algorithm based on deep learning, it has less computation and faster speed. Compared with the traditional target tracking algorithm based on Kalman filtering, it has a higher success rate and accuracy in dealing with occlusion and scale change problems.
[0169] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An event camera target tracking method based on improved Kalman filtering, characterized in that: The method includes: Step 1: accumulate the received event stream according to a fixed-length time window, and aggregate the event data in each time window into a two-dimensional event frame; Step 2: Train the support vector machine model through Gaussian radial basis function, input the two-dimensional event frame into the trained support vector machine model for occlusion judgment; if occlusion is detected, combine the feature information before occlusion in the two-dimensional event frame, use recurrent neural network to predict the target trajectory during occlusion, and after the occlusion ends, restore the position of the target according to the predicted target trajectory during occlusion; if non-occlusion is detected, directly execute step 3; Step 3: Create a Kalman filter corresponding to the target scale and initialize each Kalman filter; predict each Kalman filter based on the initialization parameters and the current position of the target to obtain the state prediction of the target at the next moment and the covariance prediction of the current state; obtain the Kalman gain based on the covariance prediction of the current state, and update each Kalman filter after the prediction through the Kalman gain to obtain the state update of the target at the next moment and the covariance update of the current state, thereby obtaining an improved Kalman filter; Step 4: Dynamically select an improved Kalman filter of corresponding scale according to the scale change of the target to perform event camera target tracking.
2. The method according to claim 1, characterized in that In step 1, the received event stream is accumulated according to a time window of a fixed length, and the event data in each time window is aggregated into a two-dimensional event frame, including: Create an event buffer, accumulate the received event stream according to a preset fixed-length time window, and use the two-dimensional array formed by all event data at each pixel position in each window as a two-dimensional event frame.
3. The method according to claim 1 or 2, characterized in that In step 2, the method of training a support vector machine model by using a Gaussian radial basis function and inputting the two-dimensional event frame into the trained support vector machine model for occlusion judgment includes: Label the dataset with occlusion labels and non-occlusion labels according to the occlusion conditions of the targets in the dataset; Use the Canny edge detection algorithm to calculate the sum of edge features of the target area in the input image or event frame; Using the sum of edge features and the annotated occlusion labels, the Gaussian radial basis function is selected to train the support vector machine model; Use the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame.
4. The method according to claim 3, characterized in that In step 2, using the Canny edge detection algorithm, the method for calculating the sum of edge features of the target area in the input image or event frame includes: 1) Perform Gaussian filtering on the input image or event frame to obtain each pixel point in the Gaussian filter template; wherein, the calculation formula for each pixel point in the Gaussian filter template is: Among them, f(x,y) is the pixel point in the Gaussian filter template, σ is the standard deviation of the Gaussian kernel, and x and y are the coordinates of a certain position of the Gaussian filter template relative to the center of the Gaussian filter template; 2) Use the Sobel operator to return the level G x Direction and vertical G y The partial derivative of the direction obtains the gradient strength G and direction θ of each pixel in the Gaussian filter template: 3) Traverse all pixels. If the gradient strength of the pixel is greater than or equal to the threshold compared with the two pixels along the gradient direction, the gradient strength is retained; if the gradient strength of the pixel is less than the threshold compared with the two pixels along the gradient direction, the gradient strength is set to 0; 4) Set a high threshold and a low threshold, mark pixels with gradient strength higher than the high threshold as edges, and pixels with gradient strength lower than the low threshold as non-edges; if the gradient strength is both greater than the low threshold and less than the high threshold, and there are edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as an edge; if there are no edge pixels with a strength greater than the high threshold among the adjacent pixels, then mark it as a non-edge; 5) Accumulate the gradient strengths of all edge pixels in the target area to obtain the sum of edge features.
5. The method according to claim 3, characterized in that In step 2, the method of using the edge feature sum and the annotated occlusion label to select the Gaussian radial basis function to train the support vector machine model includes: Use the svm module of the Scikit-learn library in Python to create a support vector machine model. Use the sum of edge features and the annotated occlusion labels, select the Gaussian radial basis function to find an optimal hyperplane or hypersurface to separate the data points of the occlusion category to be detected from the non-occlusion category, and train the support vector machine model. The Gaussian radial basis function is: Where K(m,n) is the Gaussian radial basis function value, γ is the kernel parameter, ||mn|| is the Euclidean distance calculated between the feature vectors of two samples, and m and n are the feature vectors of the two samples in the training set whose Euclidean distance is to be calculated.
6. The method according to claim 3, characterized in that In step 2, the method of using the trained support vector machine model to perform occlusion judgment on the input two-dimensional event frame includes: The two-dimensional event frame is input into the trained support vector machine model for binary classification. If the output is 0, it is judged as a non-occlusion category; if the output is 1, it is judged as an occlusion category.
7. The method according to claim 1, characterized in that In step 2, if occlusion is detected, the target trajectory during the occlusion period is predicted using a recurrent neural network in combination with feature information before the occlusion in the two-dimensional event frame. After the occlusion ends, the method for restoring the position of the target according to the predicted target trajectory during the occlusion period includes: a) Annotate the bounding box position of the target in each video frame in the dataset and divide it into training set and test set according to the preset ratio; b) Use the ResNet50 to be trained to extract image features and remove the last fully connected layer to obtain a high-level feature map; c) constructing a recurrent neural network, wherein the input of the recurrent neural network is a feature map sequence, and each time step corresponds to a feature map of a video frame; the output is the output predicted trajectory, including the position and movement speed of the target; the obtained high-level feature map is input into the recurrent neural network, and the feature map is processed by the time series to capture the temporal changes and movement patterns of the target; d) Load the training set, set the loss function to mean square error loss, and the optimizer to adaptive moment estimation optimizer; optimize the loss function through multiple rounds of iterative training, continuously adjust the temporal changes of the target and the motion pattern parameters to minimize the error between the predicted trajectory and the actual trajectory, obtain the trained recurrent neural network and test it through the test set; e) During the tracking process, when it is detected that the target is occluded, a preset number of frames of images before the occlusion are input into the trained recurrent neural network to predict the target trajectory during the occlusion period; f) Based on the predicted target trajectory, restore the target position after the occlusion ends.
8. The method according to claim 1, characterized in that In step 3, a Kalman filter corresponding to the target scale is created, and each Kalman filter is initialized including: According to whether the pixels occupied by the target scale are within the set threshold range, a Kalman filter corresponding to the threshold range is created; Initialize the state matrix, state transfer matrix, observation matrix, state covariance matrix, observation noise covariance matrix and process noise covariance matrix of each filter; where, The state matrix X represents the state of the target, including the position (x1, y1) and velocity component (v x ,v y ), expressed as: X=[x1,y1,v x ,in y ] T ; The state transfer matrix Φ is used to describe the change of state over time and is expressed as: Where Δt is the time interval; The observation matrix H is used to map the state matrix to the observation space and is expressed as: The state covariance matrix P, used to describe the uncertainty of state estimation, is expressed as: The observation noise covariance matrix R, which is used to describe the noise in the observations, is set according to the sensor noise characteristics and is expressed as: The process noise covariance matrix Q is used to describe the process noise in the state transfer process, reflecting the randomness of the target motion. It is set according to the motion characteristics of the target and is expressed as:
9. The method according to claim 1 or 8, characterized in that In step 3, the calculation method of state prediction is: X(k|k-1)=ΦX(k-1|k-1); Among them, X(k|k-1) is the system state at time k, Φ is the state transfer matrix, and X(k-1|k-1) is the optimal result of the current state; The covariance forecast is calculated as: P(k|k-1)=ΦP(k-1|k-1)Φ T +Q; Among them, P(k|k-1) is the state covariance matrix corresponding to X(k|k-1), P(k-1|k-1) is the state covariance matrix corresponding to X(k-1|k-1), and Q is the process noise covariance matrix.
10. The method according to claim 9, characterized in that In step 3, the calculation method of Kalman gain includes: Among them, K is the Kalman gain, which is used to adjust the weight between the predicted value and the current position observation value; H is the observation matrix, and R is the observation noise covariance matrix; The calculation method of the state update at the next moment includes: X(k|k)=X(k|k-1)+K(Z(k))-HX(k|k-1); Among them, X(k|k) is the state at the next moment, and Z(k) is the observation value at moment k; The calculation method of the covariance update of the current state includes: P(k|k)=(I-KH)P(k|k-1); Among them, P(k|k) is the state covariance matrix corresponding to X(k|k), and I is the identity matrix.
Citation Information
Patent Citations
Robust target tracking method based on support vector machine
CN104091349A
Anti-shielding tracking method for moving target
CN117036740A
Multi-target tracking method based on Kalman filtering and correlation matching
CN117649430A
Image tracking method based on filtering algorithm
CN118864509A
Method and apparatus for tracking moving object by using kalman filter
KR1020160110773A
Cited By
Multi-modal large model interpretation method for event stream
CN120563563A
Kalman filtering infrared target tracking method based on temperature state expansion
CN122453874A
Kalman filtering infrared target tracking method based on temperature state expansion
CN122453874B