Unsupervised Visual Object Tracking Method and System

By adopting frame skipping sampling and frame-by-frame front-back tracking training methods in the unsupervised visual target tracking method, the redundancy and inefficiency of training data are solved, and the robustness and tracking performance of the model are improved.

CN114266928BActive Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010971115.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-16
Publication Date
2025-05-30
Estimated Expiration
2040-09-16

AI Technical Summary

Technical Problem

The existing unsupervised visual target tracking methods have problems such as redundancy in training data and inefficient training, which leads to the robustness of the model being unable to meet the needs.

Method used

The frame skip sampling module is used to reduce the redundancy of training data, and data sampling is performed through jump intervals between groups and within groups, and unsupervised training is performed using frame-by-frame front-back tracking training.

Benefits of technology

It effectively reduces the amount of training data, improves training efficiency and model robustness, and improves tracking performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266928B_ABST
    Figure CN114266928B_ABST
Patent Text Reader

Abstract

An unsupervised visual object tracking method, which performs data sampling by means of inter-group jump intervals and intra-group jump intervals, takes each obtained video frame and video frame set as training samples of a siamese network architecture model for training including a frame-by-frame forward tracking process and a frame-by-frame backward tracking process, and then inputs a tracking video sequence for testing into the trained visual tracking model to obtain a finally predicted tracking box, thereby completing the tracking of the object in this frame. The present invention has good unsupervised training ability, can learn rich motion information between frames, improve training efficiency and model robustness, and performs unsupervised training through a frame-by-frame forward and backward tracking training method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of image processing, specifically an unsupervised visual object tracking method and system, which can be integrated into a similar visual object tracking Siamese network framework for efficient training, thereby improving the performance of the model, such as SiamFC. Background Art

[0002] Existing visual object tracking methods are generally trained and implemented based on supervised methods. Since supervised training requires a large amount of well-annotated training data and the tracking scenarios of the annotated data sets are subject to certain limitations, the trained models lack a certain generalization ability. Unsupervised visual object tracking methods correspondingly solve these problems, focusing on finding a suitable auxiliary task and self-supervised signal, and directly using the original data for training. Therefore, the sampled training data and the training method will substantially affect the unsupervised training of the model, and thus determine the effect of the unsupervised visual object tracking method.

[0003] Existing unsupervised correlation filtering object tracking methods often introduce the prediction task of the position index of image patches during the training of the unsupervised correlation filtering algorithm to increase the ability of the deep neural network to extract the detailed features of objects, and make the algorithm take into account both semantic information and position information by fusing the features of different layers, so as to solve the problems that the ability to extract the detailed features of objects is insufficient and it is difficult to take into account both semantic information and position information well.

[0004] However, such technologies still have the problems of redundant training data and cumbersome multi-task training processes. Summary of the Invention

[0005] In view of the deficiencies of the prior art in terms of redundant unsupervised training data, low training efficiency and the inability of the model robustness to meet the requirements, the present invention proposes an unsupervised visual object tracking method, which reduces the redundancy of training data through a frame skipping sampling module, has good unsupervised training ability, can learn rich motion information between frames, improves the training efficiency and model robustness, and performs unsupervised training through a frame-by-frame forward and backward tracking training method.

[0006] The present invention is realized through the following technical solutions:

[0007] The present invention relates to an unsupervised visual object tracking method, which performs data sampling in the way of inter-group jump interval and intra-group jump interval, takes each obtained video frame and video frame set as the training samples of the Siamese network architecture model for training including a frame-by-frame forward tracking process and a frame-by-frame backward tracking process, inputs the tracking video sequence for testing into the trained Siamese network architecture model, and obtains the finally predicted tracking box, thereby completing the tracking of the target in this frame.

[0008] The present invention relates to a system for implementing the above method, including: a frame skipping sampling module, a feature extraction module, and a correlation filter module, wherein: the frame skipping sampling module is connected to the feature extraction module and transmits training data information; the feature extraction module is connected to the correlation filter module and transmits the extracted feature information; the output of the correlation filter module is the tracking response result.

[0009] Technical effects

[0010] The present invention as a whole solves the problems of high redundancy of unsupervised training data and low training efficiency in the prior art, uses a more efficient frame-by-frame forward and backward tracking training method to improve the robustness of the model. Through the frame skipping sampling of the present invention, the amount of training data can be reduced by five times, improving the training efficiency while enhancing the tracking performance. Description of the drawings

[0011] Figure 1 It is a diagram for the implementation of the frame skipping sampling module;

[0012] Figure 2 It is a flowchart of the frame-by-frame forward and backward tracking training method. Detailed implementation manners

[0013] This embodiment relates to an unsupervised visual object tracking method, including the following steps:

[0014] Step 1) Training stage, performing data sampling, data preprocessing, and model training, specifically including:

[0015] Step 1.1) Data sampling: For a single training iteration, use the frame skipping sampling module to select a group of video frames as input in the manner of inter-group skipping interval and intra-group skipping interval.

[0016] The intra-group skipping interval is used to reduce the correlation of training data and retain valuable temporal motion information.

[0017] The inter-group skipping interval is used to ensure the full utilization of training data, reducing duplicate frames and missing frames.

[0018] As Figure 1 shown, it is a schematic diagram of the implementation of the proposed frame skipping sampling module, and the length of each video frame is 5.

[0019] Step 1.2) Data preprocessing: Perform central cropping on each video frame after frame skipping sampling, with the target size being 1 / 6 of the video frame. After padding operations, the final cropped size is 1 / 2 of the video frame, and the obtained image blocks after cropping are scaled to a size of 125x125 to obtain sets of video frames.

[0020] Step 1.3) Model Training: Use each video frame and video frame set obtained from Step 1.1 data sampling and Step 1.2 data preprocessing as training samples for the Siamese network architecture model to perform training including a frame-by-frame forward tracking process and a frame-by-frame backward tracking process.

[0021] As Figure 2 shown, the Siamese network architecture model includes a template branch and a search branch. The two branches share a feature extraction module, and this feature extraction module includes: two convolutional layers, one activation function layer, and one local response normalization layer.

[0022] In this embodiment, it is preferably set that the first convolutional layer Conv1 has a convolution kernel size of 3×3, a stride of 1×1, and 32 convolution kernels; the second convolutional layer Conv2 has a convolution kernel size of 3×3, a stride of 1×1, and 32 convolution kernels; these two convolutional layers use relatively large convolution kernels for basic feature extraction; the activation function layer is a ReLU function to alleviate the overfitting problem; the last local response normalization layer makes the values with relatively large responses become relatively larger and suppresses other values with smaller feedback, enhancing the generalization ability of the model.

[0023] The specific frame-by-frame forward tracking process includes:

[0024] First step, take the first frame of the video frame set as the template region and the second frame as the search region, and input them into the shared feature extraction module respectively to extract the template feature and the search feature; then input the template feature, the search feature, and the Gaussian response initialized at the center into the correlation filter module to obtain the search response of the second frame.

[0025] Second step, take the second frame of the video frame set as the template region and the third frame as the search region, and input them into the shared feature extraction module respectively to extract the template feature and the search feature; then input the template feature, the search feature, and the output response obtained in the first step into the correlation filter module to obtain the search response of the third frame.

[0026] Third step, repeat the first step and the second step until the search response of the fifth frame is obtained.

[0027] The specific frame-by-frame backward tracking process includes:

[0028] First step, take the fifth frame of the video frame set as the template region and the fourth frame as the search region, and input them into the shared feature extraction module respectively to extract the template feature and the search feature; then input the template feature, the search feature, and the response obtained in the last step of the frame-by-frame forward tracking process into the correlation filter module to obtain the search response of the fourth frame.

[0029] In the second step, the fourth frame of the video frame set is used as the template region, and the third frame is used as the search region, and they are respectively input into the shared feature extraction module to extract the template feature and the search feature; then the template feature, the search feature, and the output response obtained in the first step are input into the correlation filter module to obtain the search response of the third frame.

[0030] In the third step, repeat the first step and the second step until the search response of the first frame is obtained.

[0031] For the training, the loss function used is to calculate the mean square error between the Gaussian response initialized at the center and the search response obtained in the last step of the frame-by-frame backward tracking process. During training, the convolution kernel and weights are initialized with random parameters, and the bias is initialized with all zeros. The random gradient descent algorithm is used to update the model parameters. When the number of model iterations reaches the preset number of iterations, stop training and save the trained model.

[0032] Step 2) Testing stage: Input the tracking video sequence for testing into the trained Siamese network architecture model obtained in Step 1), specifically including:

[0033] 2.1) For the frame T to be tested, take the tracking box predicted in the previous frame T - 1 as the center, crop out a search image with a size of 125×125 and input it into the model. Use the incremental scale estimation scheme to handle scale changes, where each scale corresponds to a separate response map. The position of the maximum value in the response map represents the position of the tracking target, and the final predicted tracking box is obtained by combining the corresponding scale, thus completing the tracking of the target in this frame.

[0034] 2.2) Compare the tracking box predicted by the model with the tracking box annotation corresponding to the test set, and calculate the success rate and accuracy of target tracking.

[0035] In this embodiment, a Siamese network architecture model based on skip-frame sampling and trained with frame-by-frame forward and backward tracking consistency is specifically used for performance testing on the OTB - 2015 and Temple - Color - 128 datasets. Among them, the training set uses the ILSVRC2015 dataset containing 1.12 million frames as the training dataset; the test set uses the OTB - 2015 dataset containing 100 challenging sequences, including grayscale video sequences and color video sequences. The Temple - Color - 128 dataset contains 128 color sequences and poses a greater challenge.

[0036] i) For the training dataset, the final amount of training data after frame skipping sampling is 2,200 frames. Then, data preprocessing is performed on the video frames after frame skipping sampling, including central cropping and scaling, to obtain image patches of size 125×125. For a single training iteration, a set of video frame collections obtained through the above processing is represented as {I t ,I t+2 ,I t+4 ,I t+6 ,I t+8}, and used as the input for training the model.

[0037] ii) Input the training samples into the model for unsupervised training of the model, including a frame-by-frame forward tracking process and a frame-by-frame backward tracking process, where:

[0038] The frame-by-frame forward tracking process includes:

[0039] In the first step, in this embodiment, I t is used as the template region, and I t+2 is used as the search region, and they are respectively input into the shared feature extraction module to extract the template feature T t and the search feature T t+2 . The template feature T t , the search feature T t+2 , and the Gaussian response Y t initialized at the center are input into the correlation filter module to obtain the search response R t+2 of I t,t+2 .

[0040] In the second step, in this embodiment, I t+2 is used as the template region, and I t+4 is used as the search region, and they are respectively input into the shared feature extraction module to extract the template feature T t+2 and the search feature T t+4 . The template feature T t+2 , the search feature T t+4 , and the output response R t,t+2 from the previous step are input into the correlation filter module to obtain the search response R t+4 of I t+2,t+4 .

[0041] In the third step, in this embodiment, I t+4 is used as the template region, and I t+6 is used as the search region, and they are respectively input into the shared feature extraction module to extract the template feature T t+4 and the search feature T t+6 . The template feature T t+4 , the search feature T t+6 , and the output response R t+2,t+4Input it into the correlation filter module to obtain the search response R of I t+6 of I t+4,t+6 .

[0042] In the fourth step, this embodiment uses I t+6 as the template area and I t+8 as the search area, and inputs them into the shared feature extraction module respectively to extract the template feature T t+6 and the search feature T t+8 . Input the template feature T t+6 , the search feature T t+8 and the output response R of the previous step t+4,t+6 into the correlation filter module to obtain the search response R of I t+8 of I t+6,t+8 .

[0043] The frame-by-frame backward tracking process includes:

[0044] In the first step, this embodiment uses I t+8 as the template area and I t+6 as the search area, and inputs them into the shared feature extraction module respectively to extract the template feature T t+8 and the search feature T t+6 . Input the template feature T t+8 , the search feature T t+6 and the output response R of the last step of forward tracking t+8,t+6 into the correlation filter module to obtain the search response R of I t+6 of I t+8,t+6 .

[0045] In the second step, this embodiment uses I t+6 as the template area and I t+4 as the search area, and inputs them into the shared feature extraction module respectively to extract the template feature T t+6 and the search feature T t+4 . Input the template feature T t+62 , the search feature T t+4 and the output response R of the previous step t+8,t+6 into the correlation filter module to obtain the search response R of I t+4 of I t+6,t+4 .

[0046] In the third step, this embodiment uses I t+4 as the template area and I t+2 as the search area, and inputs them into the shared feature extraction module respectively to extract the template feature T t+4 and the search feature T t+2 . Input the template feature T t+4 , the search feature T t+2and the output response R of the previous step t+6,t+4 is input into the correlation filter module to obtain the search response R t+2 of I t+4,t+2 .

[0047] In the fourth step, in this embodiment, I t+2 is used as the template region, and I t is used as the search region, and they are respectively input into the shared feature extraction module to extract the template feature T t+2 and the search feature T t . The template feature T t+2 , the search feature T t and the output response R of the previous step t+4,t+2 are input into the correlation filter module to obtain the search response R t of I t+2,t .

[0048] The loss function for training is to calculate the mean square error between the Gaussian response Y t initialized at the center and the search response R t+2,t . During the training process, the convolutional kernels and weights in the shared feature extraction module are randomly initialized, and the bias terms are set to 0. The random gradient descent algorithm is used to update the model parameters. When the number of model iterations reaches the preset value, the training stops and the trained model is saved.

[0049] Table 1 Parameter settings of the shared feature extraction module

[0050]

[0051] iii) Input the test tracking video sequence into the trained visual tracking model, compare the tracking box predicted by the model with the tracking box annotation corresponding to the test set, and calculate the success rate and accuracy of object tracking. Among them, the success rate is the proportion of the overlap rate between the predicted tracking box and the annotated tracking box being greater than a given threshold. The accuracy is the proportion of the distance between the center point of the predicted tracking box and the center point of the annotated tracking box within different distance pixel ranges.

[0052] As shown in Table 2 and Table 3, the method of this embodiment can achieve good results on different public datasets and obtains the best results among all unsupervised visual object tracking methods.

[0053] Table 2 Performance comparison of different visual object tracking methods on the OTB-2015 dataset

[0054]

[0055] Table 3 Performance comparison of different visual object tracking methods on the Temple-Color-128 dataset

[0056]

[0057] Compared with the existing sampling techniques that randomly sample some frames from a video sequence to train unsupervised representations, rich motion information will be discarded. Moreover, such a set of randomly sampled frames may include duplicate frames, resulting in data redundancy. For a single training iteration, the present invention selects a set of video frames as input in the way of intra-group jump interval and inter-group jump interval through frame skipping sampling, where the intra-group jump interval is used to reduce the correlation of training data and retain valuable temporal motion information, and the inter-group jump interval is used to ensure the full utilization of training data, reduce duplicate frames and missing frames and achieve a good balance between reducing the correlation of training data and retaining valuable motion information between video sequences.

[0058] The above specific implementation can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation, and all implementation solutions within its scope are subject to the present invention.

Claims

1. An unsupervised visual object tracking method, characterized in that, data sampling is performed by means of inter-group jump intervals and intra-group jump intervals, and each obtained video frame and video frame set are used as training samples for a siamese network architecture model for training including a frame-by-frame forward tracking process and a frame-by-frame backward tracking process. The tracking video sequence for testing is input into the trained siamese network architecture model to obtain the finally predicted tracking box, thereby completing the tracking of the object in this frame; the siamese network architecture model includes a template branch and a search branch, and the two branches share a feature extraction module, and the feature extraction module includes: two convolutional layers, an activation function layer and a local response normalization layer; the specific frame-by-frame forward tracking process includes: First step, the first frame of the video frame set is used as the template area and the second frame is used as the search area, and they are respectively input into the shared feature extraction module to extract the template feature and the search feature; then the template feature, the search feature and the Gaussian response initialized at the center are input into the correlation filter module to obtain the search response of the second frame; Second step, the second frame of the video frame set is used as the template area and the third frame is used as the search area, and they are respectively input into the shared feature extraction module to extract the template feature and the search feature; then the template feature, the search feature and the output response obtained in the first step are input into the correlation filter module to obtain the search response of the third frame; Third step, repeat the first step and the second step until the search response of the fifth frame is obtained; the specific frame-by-frame backward tracking process includes: First step, the fifth frame of the video frame set is used as the template area and the fourth frame is used as the search area, and they are respectively input into the shared feature extraction module to extract the template feature and the search feature; then the template feature, the search feature and the response obtained in the last step of the frame-by-frame forward tracking process are input into the correlation filter module to obtain the search response of the fourth frame; Second step, the fourth frame of the video frame set is used as the template area and the third frame is used as the search area, and they are respectively input into the shared feature extraction module to extract the template feature and the search feature; then the template feature, the search feature and the output response obtained in the first step are input into the correlation filter module to obtain the search response of the third frame; Third step, repeat the first step and the second step until the search response of the first frame is obtained.

2. The method according to claim 1, characterized in that for the data sampling, for a single training iteration, a frame skipping sampling module is used to select a set of video frames as input in the manner of inter-group jump intervals and intra-group jump intervals, where the intra-group jump interval is used to reduce the correlation of training data and retain valuable temporal motion information; the inter-group jump interval is used to ensure the full utilization of training data and reduce duplicate frames and missing frames.

3. The method according to claim 1 or 2, characterized in that the intra-group jump interval is 2 and the inter-group jump interval is 5.

4. The method according to claim 1, characterized in that The training samples are obtained by preprocessing the sampled data, specifically: each video frame after skipping-frame sampling is centrally cropped, with the target size being 1 / 6 of the video frame. After a padding operation, the final cropped size is 1 / 2 of the video frame, and the obtained image patches are scaled to a size of 125x125 to obtain sets of video frame collections.

5. The method according to claim 1, characterized in that for the training, the loss function used is to calculate the mean square error between the Gaussian response initialized at the center and the search response obtained in the last step of the frame-by-frame backward tracking process. During training, the convolution kernel and weights are initialized with random parameters, and the bias is initialized with all zeros.

6. The method according to claim 1, characterized in that for the training, the random gradient descent algorithm is used to update the model parameters. When the number of model iterations reaches the preset number of iterations, the training is stopped and the trained model is saved.

7. A system for implementing the method according to any one of claims 1-6, characterized in that it includes: a skipping-frame sampling module, a feature extraction module, and a correlation filter module, where: the skipping-frame sampling module is connected to the feature extraction module and transmits training data information; the feature extraction module is connected to the correlation filter module and transmits the extracted feature information; the output of the correlation filter module is the tracking response result.

Citation Information

Patent Citations

  • Target specific response attention target tracking method based on twin network

    CN111291679A

  • Target tracking method and device oriented to airborne-based monitoring scenarios

    US20200051250A1