A Visual Single Object Tracking Method Assisted by Motion Information
By introducing camera motion and target motion detectors and validators into the target tracking algorithm, and calculating and fusing motion offsets to correct the search area, the problem of loss of appearance information or reduced robustness under occlusion in the prior art is solved, and the robustness and accuracy of the tracking algorithm are improved.
Patent Information
- Application Number
- CN202211532172.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing target tracking algorithms are less robust when appearance information is missing or obstructed, making it difficult to effectively cope with the influence of factors such as camera movement and lighting changes.
A motion information-assisted visual single-object tracking method is proposed. The target offset caused by camera motion is calculated by the camera motion detector and the validator, and the target motion offset is estimated by the target motion detector and the validator, combining the two to correct the search area of the visual tracker.
It effectively solves the problem of target removal search area caused by camera movement, improves the robustness of the tracking algorithm, and reduces the risk of losing targets caused by occlusion and deformation.
Smart Images

Figure CN116309683B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning, object tracking, and object trajectory prediction, and relates to image feature point matching, the single-object tracking algorithm DiMP, and the object feature extraction network ResNet. Specifically, the present invention relates to a visual single-object tracking method assisted by motion information. Background Art
[0002] In the field of object tracking, object appearance information is often affected by factors such as occlusion, non-rigid deformation, and illumination changes. Although existing methods use efficient feature extraction networks to model object appearance information, the robustness of the tracker is significantly reduced in scenarios where appearance information is missing. Among them, motion information is insensitive to scene and object changes, so it can be used as auxiliary information to improve tracking accuracy and robustness. Based on the above content, the present invention mainly explores the potential of motion information in single-object tracking and conducts further analysis based on this.
[0003] Single-object tracking algorithms output the position and scale information of the target in subsequent frames by given the position of the target in the first frame of the video. Early target tracking algorithms include Kalman filtering, particle filtering, etc.; these methods locate the current target position through the target's historical trajectory and velocity information. In recent years, with the emergence of deeper networks with stronger discriminative ability, visual trackers based on appearance information have achieved high accuracy. Commonly used visual trackers based on appearance information include correlation filter algorithms (correlation filter [Henriques, J.F., Caseiro, R., Martins, P., & Batista, J. High-speed tracking with kernelized correlation filters.]), Siamese networks (Siamese network [Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., & Torr, P.H.F. Fully-convolutional siamese networks for object tracking.]), and multi-domain networks (multi-domain network [Nam, H., & Han, B. Learning multi-domain convolutional neural networks for visual tracking.]). Since motion information is insensitive to scene changes, motion information can help appearance information improve the robustness of the algorithm and better cope with factors such as occlusion and illumination changes. Existing algorithms improve tracking accuracy and robustness by modeling the camera and using spatio-temporal networks to predict the target trajectory, and explore the role of motion information in the tracking scene. Summary of the Invention
[0004] The present invention aims to propose an integrated motion information modeling framework to simultaneously model the target motion information and the camera motion information.
[0005] The method of the present invention can be deployed on edge devices such as drones and security surveillance cameras to act as a visual module to achieve target tracking.
[0006] The technical solution of the present invention is as follows:
[0007] A visual single-object tracking method assisted by motion information, the steps are as follows:
[0008] Step 1: Obtain continuous video frames of the tracking area with the help of an airborne camera of a drone or a surveillance camera;
[0009] Step 2: Input a continuous video stream. According to the timing relationship between the previous frame and the current frame, calculate the target offset caused by camera movement and the corresponding confidence level for the current frame through a camera movement detector and a camera movement validator respectively;
[0010] Operation process of the camera movement detector: Extract the SIFT feature points of the video frame, use the RANSAC method to match the SIFT feature points of the video frame, and calculate the affine transformation matrix M between the images through formula (1):
[0011]
[0012] where, It includes feature point extraction and image matching operations; where I t-1 represents the image at time t - 1, and I t represents the image at time t, that is, the current frame image; Finally, apply the obtained affine transformation matrix M to the target rectangular box to obtain the target offset O caused by camera movement CM :
[0013]
[0014] O CM = c′ - c
[0015] where, c represents the target center point coordinates on the I t-1 frame, and c′ represents the target center point coordinates of the current frame estimated by the affine transformation matrix;
[0016] Operation process of the camera movement validator: The camera movement validator includes 5 convolutional layers with a convolutional kernel of 3 * 3 and 3 fully connected layers; Obtain the transformed image of the search area of the current frame visual tracker from the affine transformation matrix, and use the current frame image, the transformed image, and their corresponding video frame SIFT feature points as the input of the camera movement validator, and output the confidence level S of the current camera transformation CM ;
[0017] Step 3: According to the target historical trajectory, estimate the target movement offset and the corresponding confidence level for the current frame through a target movement detector and a target movement validator;
[0018] Given the target historical trajectory, that is, the target historical position and speed The target movement detector is expected to give the current position of the target and speed The input of the target movement detector is the motion heat map G t generated by x t , where G t obeys a two-dimensional Gaussian distribution Where g is the width and height of the heat map, and W and H are the width and height of the current frame respectively; the target motion detector includes two convolutional layers with a convolutional kernel of 5*5, an LSTM network, and two fully connected layers; the motion heat map G t First, it passes through two convolutional layers and is fed into the LSTM network, and finally, the normalized motion vector x is output by two fully connected layers t+1 , and finally, the motion heat map of the current frame is obtained according to the normalized motion vector Finally, the target motion offset of the current frame is obtained through formula (3):
[0019]
[0020] And based on the offset and the position of the target in the previous frame, the position of the target in the current frame is predicted
[0021] The target motion validator includes three convolutional layers, one region of interest pooling layer, and three fully connected layers; the target motion validator takes the image at the predicted position of the target in the current frame as input, determines whether the current target information is accurate, and outputs the confidence S of the current target motion information OM ;
[0022] Step 4: Fuse the motion offsets and confidences obtained in Step 2 and Step 3 to obtain the total offset O:
[0023] O = r(S CM ) * O CM + r(S OM ) * O OM (4) Where r() represents the step function;
[0024]
[0025] Use the obtained offset to correct the search area position of the visual tracker
[0026] Step 5: Apply the offset calculated using the motion information to the visual tracker to achieve robust tracking; the visual tracker includes a ResNet feature extractor, an IoU prediction network, and an online classifier, and finally outputs the target position response map of the current frame to obtain the final tracking result; first, send the corrected search area of the current frame into the ResNet feature extractor to extract image features; send the obtained image features into the online classifier and the IoU prediction network respectively; for the online classifier, use the learned model to give the target position response map and update the model using the information of the current frame, and for the IoU prediction network, the network uses the IoUNet structure to predict the target scale change
[0027] The beneficial effects of the present invention:
[0028] (1) The algorithm of the present invention models the camera movement, calculates the corresponding transformation matrix of adjacent frames by using the feature point matching method, and corrects the search area of the algorithm, effectively solving the problem of the target removal search area caused by the camera movement.
[0029] (2) The method of the present invention uses a convolutional long short-term memory network to model the target trajectory and speed, and predicts the possible positions of the target, reducing the risk of losing the target caused by target occlusion and deformation, and improving the robustness of the tracking algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is the overall flowchart of the present invention.
[0031] Figure 2 is the structural schematic diagram of the camera motion detector and validator.
[0032] Figure 3 is the structural flowchart of the target motion detector and validator. DETAILED DESCRIPTION OF THE INVENTION
[0033] The following further describes the specific implementation manners of the present invention in conjunction with the drawings and technical solutions.
[0034] Figure 1 is the overall flowchart of the present invention. The algorithm consists of three parts, namely, the camera motion detector and validator, the target motion detector and validator, and the visual single-target tracker. The camera motion detector and validator are responsible for detecting the image changes caused by the camera movement through the structural relationship of adjacent frames, calculating the target box offset, and giving the confidence of the current estimated offset. The structure of the camera motion detector and validator is as Figure 2 shown. First, the SIFT feature points of the current frame and the historical frame images are extracted and matched to calculate the camera motion offset and the affine transformation matrix. Then, the transformed historical frame image and the current frame image are sent into a convolutional network composed of 5 convolutional layers, and the features of the two images are concatenated and sent into a fully connected network to obtain the camera motion confidence. The fully connected network contains 3 fully connected layers.
[0035] Figure 3It is the structural flow chart of the target motion detector and validator. First, the target historical information is position-encoded through a position encoder to obtain a motion heat map, and a single-layer convolutional network is used to extract features. The temporal heat map data is input into a long short-term memory network, and a normalized motion vector is obtained through a fully connected network with two fully connected layers. Finally, the current frame's motion heat map is decoded from the normalized motion vector to achieve position prediction. The camera motion validator first crops the predicted position image and uses the obtained image as the input of a convolutional network for feature extraction, where the convolutional network consists of three convolutional layers. The extracted features are fed into a region of interest pooling layer to obtain a region of interest vector, and the target motion confidence is estimated by inputting it into a network with three fully connected layers.
[0036] After predicting the camera motion and target motion offsets respectively, we aggregate the two types of motion information through their respective motion confidences and correct the search region of the single-object tracker. The present invention can be applied to most visual single-object trackers to implement the embedding of motion information. Figure 1 Taking the DiMP tracking algorithm as an example, the method uses ResNet as a feature extractor to extract features from the current frame's search region. The extracted features are respectively used to estimate the target scale and target position. For target scale estimation, the tracker uses IoUNet as an IoU prediction network. For target position estimation, DiMP adopts an online-updatable convolutional layer as an online classifier to model the temporal changes of the target.
[0037] The method of the present invention is trained using the LaSOT dataset. The training data of the camera motion detector and validator realizes camera motion estimation by transforming the LaSOT dataset. The present invention uses the Adam optimization method for learning, with an initial learning rate of 0.0001, and adopts a learning rate decay strategy for learning, and the number of training epochs is 300. Among them, the hyperparameter t in motion aggregation is 0.6.
Claims
1. A visual single-object tracking method assisted by motion information, characterized in that The steps are as follows: Step 1: Obtain consecutive video frames of the tracking area with the help of an airborne drone camera or a surveillance camera; Step 2: Input the consecutive video stream. According to the temporal relationship between the previous frame and the current frame, calculate the target offset caused by camera movement and the corresponding confidence level in the current frame through a camera movement detector and a camera movement validator respectively; Operation process of the camera movement detector: Extract SIFT feature points of the video frame, use the RANSAC method to match the SIFT feature points of the video frame, and calculate the affine transformation matrix M between images through formula (1): Among them, including feature point extraction and image matching operations; where I t-1 represents the image at time t-1, I t represents the image at time t, that is, the current frame image; finally, the obtained affine transformation matrix M is applied to the target rectangular box to obtain the target offset O caused by camera movement CM : where c represents the coordinates of the target center point on the I t-1 frame, and c' represents the coordinates of the target center point of the current frame estimated by the affine transformation matrix; Operation process of the camera motion validator: The camera motion validator includes 5 convolutional layers with a convolutional kernel of 3*3 and 3 fully connected layers; the transformed image obtained from the affine transformation matrix of the search area of the current frame visual tracker is used, and the current frame image, the transformed image, and their corresponding video frame SIFT feature points are used as the input of the camera motion validator, and the confidence S of the current camera transformation is output CM ; Step 3: According to the target historical trajectory, estimate the target movement offset and the corresponding confidence level in the current frame through a target movement detector and a target movement validator; Given the target historical trajectory, i.e., the target historical position and velocity The target motion detector is expected to give the target current position and velocity The input of the target motion detector is the motion heat map G t generated by x t , where G t follows a two-dimensional Gaussian distribution where g is the width and height of the heat map, and W and H are the width and height of the current frame respectively; the target motion detector includes two convolutional layers with a convolutional kernel of 5*5, an LSTM network and two fully connected layers; the motion heat map G t First, it passes through two convolutional layers and is fed into the LSTM network, and finally, the normalized motion vector x t+1 is output by two fully connected layers, and finally, the motion heat map of the current frame is obtained according to the normalized motion vector Finally, the target motion offset of the current frame is obtained through formula (3): Predict the position of the target in the current frame based on the offset and the position of the target in the previous frame; The target motion validator includes three convolutional layers, one region of interest pooling layer, and three fully connected layers; the target motion validator takes the image at the position where the target is predicted to be in the current frame as input, determines whether the current target information is accurate, and outputs the confidence level S of the current target motion information OM ; Step 4: Fuse the movement offsets and confidence levels obtained in Step 2 and Step 3 to obtain the total offset O: O = r(S CM ) * O CM + r(S OM ) * O OM (4) where r() represents the step function; Use the obtained offset to correct the search area position of the visual tracker; Step 5: Apply the offset calculated using motion information to the visual tracker to achieve robust tracking; The visual tracker includes a ResNet feature extractor, an IoU prediction network, and an online classifier, and finally outputs the target position response map of the current frame to obtain the final tracking result; First, send the corrected search area of the current frame into the ResNet feature extractor to extract image features; Send the obtained image features into the online classifier and the IoU prediction network respectively; For the online classifier, use the learned model to give the target position response map and update the model with the information of the current frame. For the IoU prediction network, the network uses the IoUNet structure to predict the target scale change.
Citation Information
Patent Citations
Rapid movement compensation method of moving target detection under mobile camera
CN106534614A
An online multi-target tracking method based on R-FCN framework multi-candidate association
CN109919974A
Cited By
Unmanned aerial vehicle visual perception processing method, system and device based on motion event optical flow and medium
CN121937490A