Visible and Infrared Visual Tracking Method Based on Meta-Learning Parameter Transfer
By introducing meta-learning parameter transfer technology into the multimodal visual tracking control model, the problem of insufficient data volume in multimodal target tracking is solved, and more efficient and accurate target tracking performance is achieved.
Patent Information
- Application Number
- CN202111625448.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-28
AI Technical Summary
In the prior art, the multimodal target tracking originates too late, requiring pair-matched information, high collection cost and limited data volume, which limits performance improvement.
Using the visible light infrared vision tracking method based on meta-learning parameter transfer, by constructing a multimodal vision tracking control model, first input image pairs into the model for iterative training, and then alternately input separate visible light images and image pairs, and introduce a meta-learner to make up for the problem of insufficient data.
It effectively improves the performance of multimodal visual tracking, solves the problem of limited model performance improvement caused by limited data volume, and achieves more accurate and efficient target tracking.
Smart Images

Figure CN114299114B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and more particularly to a visible light and infrared vision tracking method based on meta-learning parameter transfer. Background Art
[0002] Visual object tracking is an important research direction in the field of computer vision. Object tracking has a wide range of application fields in the industrial community, such as security monitoring, autonomous driving, behavior analysis, etc.
[0003] Currently, object tracking based on the detection and tracking framework is essentially a binary classification problem of the object and the background based on a well-performing classifier. For example, in the master's thesis "Research on Object Tracking Algorithm and Analysis Based on Binary Classifier" published by Zhang Bobin of China University of Mining and Technology, object tracking is carried out based on the particle filter framework. Traditional object tracking first determines the cropped samples and determines whether they are the object or the background by setting a threshold and comparing the coincidence rate between the object and the sample with the threshold at the given object position in the first frame, and initializes the object tracking model with these samples, and then performs the tracking of the object in the subsequent frames. In the subsequent frames, Gaussian distribution sampling is still performed near the predicted object position in the previous frame, and a classifier is used to distinguish between object samples and background samples, and the mean value of the top five highest-scoring positive samples is selected as the initial position of the object in the next frame.
[0004] The modalities used in multi-modal object tracking are the visible light modality and the thermal infrared modality. Visible light (wavelength 0.4 - 0.7μm) images contain rich geometric and texture details and can well distinguish the edge information of adjacent objects, but are sensitive to light changes and bad weather. Thermal infrared (wavelength 3 - 12μm) images reflect the surface temperature distribution in the scene and effectively resist the influence brought by light changes and bad weather conditions. The complementarity of these two modalities further improves the performance of object tracking. However, the limited amount of multi-modal data restricts the performance improvement. For example, the performance tested using the public dataset RGBT234 is always much lower than that of the public dataset GTOT. Part of the reason is that the amount of multi-modal data used for training is relatively small, and the only multi-modal data limits the performance expression of the network. The origin of multi-modal tracking is relatively late and requires paired matching information. The acquisition of these paired datasets requires special cameras and manual alignment, and the labor cost is expensive, resulting in a huge difference in the amount of multi-modal data and single-modal data. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the origin of existing multi-modal object tracking is relatively late and requires paired matching information. The acquisition of these paired data sets requires special cameras and manual alignment, which is costly in terms of labor and has limited data volume. The limited multi-modal data volume restricts the improvement of performance.
[0006] The present invention solves the above technical problem by the following technical means: a visible light and infrared vision tracking method based on meta-learning parameter transfer, the method comprising:
[0007] Step 1: Construct a multi-modal vision tracking control model;
[0008] Step 2: Input the samples into the multi-modal vision tracking control model for training. The samples include multiple pairs of thermal infrared images and their corresponding visible light images, as well as multiple individual visible light images. During the training process, first input the image pairs into the model for the first preset number of iterative trainings, and then alternately input the individual visible light images and the image pairs to continue training the model for the second preset number of times. And when using the image pairs for model training during the alternating training process, the images need to pass through a meta-learner;
[0009] Step 3: After the training is completed, collect videos in real time, extract paired thermal infrared images and their corresponding visible light images, input them into the trained model, and the model tracks and outputs the predicted target position.
[0010] During the process of training the multi-modal vision tracking control model of the present invention, first input the image pairs into the model for the first preset number of iterative trainings, and then alternately input the individual visible light images and the image pairs to continue training the model for the second preset number of times. The paired image data set is relatively small, and additional visible light images are added to participate in the training to make up for the deficiency of the training set. At the same time, a meta-learner is introduced to further solve the situation of insufficient network training caused by insufficient data sets in the prior art, and to avoid the model from converging too quickly or not converging due to a small data set, so as to avoid inaccurate training results and improve the model performance, and solve the problem that the performance improvement of the model is restricted due to limited data volume in the prior art.
[0011] Furthermore, the multi-modal visual tracking control model includes a first general adapter to a third general adapter, a first visible light modal adapter to a third visible light modal adapter, a first thermal infrared modal adapter to a third thermal infrared modal adapter, a meta-learner, a first instance adapter, and a second instance adapter, which are sequentially numbered. The first visible light modal adapter is connected to an input end of the first general adapter and receives the visible light image corresponding to a single visible light image or thermal infrared image. The first thermal infrared modal adapter is connected to the other input end of the first general adapter and receives the thermal infrared image. The output result of the first visible light modal adapter is superimposed on the first output end of the first general adapter and then input into the second visible light modal adapter and the second general adapter. The output result of the first thermal infrared modal adapter is superimposed on the second output end of the first general adapter and then input into the second thermal infrared modal adapter and the second general adapter. The output result of the second visible light modal adapter is superimposed on the first output end of the second general adapter and then input into the third visible light modal adapter and the third general adapter. The output result of the second thermal infrared modal adapter is superimposed on the second output end of the second general adapter and then input into the third thermal infrared modal adapter and the third general adapter. The output result of the third visible light modal adapter is superimposed on the first output end of the third general adapter. The third visible light modal adapter is connected to the third thermal infrared modal adapter through the meta-learner. The output result of the third thermal infrared modal adapter is superimposed on the second output end of the third general adapter through the dimensionality reduction unit. When only the output result of the first output end of the third general adapter is available, this output result is output through the second instance adapter. Otherwise, the results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
[0012] Furthermore, the first visible light modality adapter and the first thermal infrared modality adapter have the same structure, both consisting of a 3×3 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the second visible light modality adapter and the second thermal infrared modality adapter have the same structure, both consisting of a 1×1 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the third visible light modality adapter, the third thermal infrared modality adapter, and the dimensionality reduction unit have the same structure, both consisting of a 1×1 convolutional layer, ReLU, and LRN cascaded in sequence; the first general adapter consists of a 7×7 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the second general adapter consists of a 5×5 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the third general adapter consists of a 3×3 convolutional layer, ReLU, and LRN; the meta-learner consists of two learning units cascaded in sequence, both learning units consisting of ReLU and a fully connected layer, and the dimensionalities of the fully connected layers of the two learning units are different; the first instance adapter consists of a fully connected layer FC4, a ReLU, a dropout function, a fully connected layer FC5, another ReLU, another dropout function, and a fully connected layer FC6 cascaded in sequence, and the second instance adapter has the same structure as the first instance adapter, but the dimensionality of the fully connected layer FC4 is different.
[0013] Furthermore, the training process of the multi-modal visual tracking control model includes:
[0014] Iteratively pre-train the network model 70 times using multiple image pairs formed by thermal infrared images and their corresponding visible light images, set the batch size to 128, the learning rate of the convolutional layer to 0.0001, and the learning rate of the fully connected layer to 0.0002. During the training process, the visible light image and its corresponding thermal infrared image are input into the model simultaneously. The visible light image passes through the first visible light adapter and the corresponding first general adapter, the second visible light adapter and the corresponding second general adapter, the third visible light adapter and the corresponding third general adapter in sequence, and the output result is fused to the first output end of the third general adapter. The thermal infrared image passes through the first thermal infrared adapter and the corresponding first general adapter, the second thermal infrared adapter and the corresponding second general adapter, the third thermal infrared adapter and the corresponding third general adapter in sequence, and the output result is fused to the second output end of the third general adapter. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
[0015] Furthermore, the sample acquisition method used during the training process is: in each frame of the video, select S + = 4 (IOU≥0.7) and S - = 12 (IOU≤0.5) sample numbers, where S+ Indicates a positive sample, S - Indicates a negative sample. IOU represents the intersection over union between the collected sample and the ground truth box, through the collected positive and negative samples.
[0016] Furthermore, the training process of the multi-modal visual tracking control model further includes: alternately inputting a single visible light image and an image pair to continue training the model 130 times. The learning rate and the sample sampling method are the same as those used for iteratively pre-training the network model with 70 times of a thermal infrared image and its corresponding visible light image pair.
[0017] Furthermore, each time a single visible light image is used to train the model, the single visible light image passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position in sequence, and then the output result is fused to the first output end of the third general adapter. The output result of the first output end of the third general adapter is output from the second instance adapter.
[0018] Furthermore, during the process of alternately inputting a single visible light image and an image pair to continue training the model 130 times, each time the image pair is used to train the model, the visible light image passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position in sequence, and then the output result is fused to the first output end of the third general adapter. The thermal infrared image passes through the first thermal infrared adapter and the first general adapter at the corresponding position, the second thermal infrared adapter and the second general adapter at the corresponding position in sequence. The output results of the second thermal infrared adapter and the second general adapter at the corresponding position are fused and then input to the third thermal infrared adapter and the third general adapter. The output of the third visible light adapter is input to the third thermal infrared adapter through the meta-learner. The output result of the third thermal infrared adapter is fused to the second output end of the third general adapter after passing through the dimensionality reduction unit. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
[0019] Further, the third step includes: After the training is completed, videos are collected in real time, and paired thermal infrared images and their corresponding visible light images are extracted and input into the trained model. The visible light images pass through the first visible light adapter and the corresponding first general adapter, the second visible light adapter and the corresponding second general adapter, and the third visible light adapter and the corresponding third general adapter in sequence, and then the output results are fused to the first output end of the third general adapter. The thermal infrared images pass through the first thermal infrared adapter and the corresponding first general adapter, the second thermal infrared adapter and the corresponding second general adapter, and the third thermal infrared adapter and the corresponding third general adapter in sequence, and then the output results are fused to the second output end of the third general adapter. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter to output the predicted target position.
[0020] Furthermore, the output through the first instance adapter to output the predicted target position includes: The fully connected layer FC6 of the first instance adapter contains a softmax layer to calculate the positive and negative scores for each sample feature: f + (x i ) and f - (x i ). The predicted target position is obtained through the formula , where x i represents the i-th sampled sample, f + (x i ) represents the obtained positive sample score, f - (x i ) represents the obtained negative sample score, and x * is the predicted target position.
[0021] The advantages of the present invention are as follows:
[0022] (1) During the process of training the multi-modal visual tracking control model of the present invention, image pairs are first input into the model for the first preset number of iterative trainings, and then single visible light images and image pairs are alternately input to continue training the model for the second preset number of times. The paired image dataset is relatively small, and additional visible light images are added to participate in the training to make up for the deficiency of the training set. At the same time, a meta-learner is introduced to further solve the situation of insufficient network training caused by insufficient dataset in the prior art, avoid the model from converging too quickly or being unable to converge due to a small dataset, thereby avoiding inaccurate training results and improving the model performance, and solving the problem that the model performance improvement is limited due to limited data volume in the prior art.
[0023] (2) The present invention first uses multiple image pairs formed by thermal infrared images and their corresponding visible light images to iteratively train the network model 70 times, without passing through a meta-learner or a dimensionality reduction unit, and then outputs from the first instance adapter. First, the main framework of the model is trained to a better state, avoiding the slow training process and inaccurate model results caused by extracting some unimportant features when directly training all units of the model with images.
[0024] (3) The present invention alternately inputs individual visible light images and image pairs to continue training the model 130 times, solving the problem of relatively few existing paired image datasets, adding additional visible light images to participate in training, and making up for the deficiencies of the training set.
[0025] (4) The present invention utilizes existing unimodal data to improve the performance of multimodal tracking, and proposes a meta-learner that can transfer the learning ability of the RGB modality to the T modality in paired multimodal information, because at this time, the RGB modality has stronger learning ability after being trained on more datasets.
[0026] (5) The present invention enhances the feature learning ability of the visible light modality adapter by introducing additional unimodal data, and this learning ability is transferred to the thermal infrared modality adapter through the introduced meta-learner by parameter transfer. That is, it makes full use of the additional unimodal data and does not need to generate its corresponding thermal infrared dataset by itself, reducing the input of noise. The meta-learner can be removed during the tracking process to obtain better tracking performance without increasing additional testing time. Description of the Drawings
[0027] Figure 1 It is a model architecture diagram of the visible light-infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention;
[0028] Figure 2 It is a flow chart of the tracking process of the visible light-infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention;
[0029] Figure 3 It is a comparison chart of the success rates of the visible light-infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention and other trackers tested on the GTOT dataset;
[0030] Figure 4 It is a comparison chart of the precisions of the visible light-infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention and other trackers tested on the GTOT dataset;
[0031] Figure 5Comparison chart of success rates of the visible-light infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention and other trackers tested on the RGBT234 dataset;
[0032] Figure 6 Comparison chart of precisions of the visible-light infrared visual tracking method based on meta-learning parameter transfer disclosed in the embodiments of the present invention and other trackers tested on the RGBT234 dataset. Detailed implementation manners
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] A visible-light infrared visual tracking method based on meta-learning parameter transfer, the method comprising:
[0035] Step 1: Construct a multi-modal visual tracking control model;
[0036] Step 2: Input the samples into the multi-modal visual tracking control model for training. The samples include multiple pairs of thermal infrared images and their corresponding visible-light images, as well as multiple individual visible-light images. During the training process, first input the image pairs into the model for the first preset number of iterative trainings, and then alternately input the individual visible-light images and the image pairs to continue training the model for the second preset number of times. And when using the image pairs for model training during the alternating training process, the images need to pass through a meta-learner; wherein, the acquisition method of the training samples is: acquire the first frame of a pair of registered multi-modal videos, and through the given ground truth box of the target in the first frame, perform Gaussian distribution sampling with the center point of the ground truth box as the mean and (0.09r2, 0.09r2, 0.25) as the covariance to generate 256 candidate samples, where: r is the average value of the sum of the width and height of the target in the previous frame.
[0037] Step 3: After the training is completed, collect videos in real time, extract paired thermal infrared images and their corresponding visible-light images, input them into the trained model, and the model tracks and outputs the predicted target position.
[0038] As Figure 1As shown, the multimodal visual tracking control model includes a first general adapter to a third general adapter, a first visible light modality adapter to a third visible light modality adapter, a first thermal infrared modality adapter to a third thermal infrared modality adapter, a meta-learner, a first instance adapter, and a second instance adapter, which are sequentially numbered. The first visible light modality adapter is connected to an input end of the first general adapter and receives a visible light image corresponding to a single visible light image or a thermal infrared image. The first thermal infrared modality adapter is connected to the other input end of the first general adapter and receives a thermal infrared image. The output result of the first visible light modality adapter is superimposed on the first output end of the first general adapter and then input into the second visible light modality adapter and the second general adapter. The output result of the first thermal infrared modality adapter is superimposed on the second output end of the first general adapter and then input into the second thermal infrared modality adapter and the second general adapter. The output result of the second visible light modality adapter is superimposed on the first output end of the second general adapter and then input into the third visible light modality adapter and the third general adapter. The output result of the second thermal infrared modality adapter is superimposed on the second output end of the second general adapter and then input into the third thermal infrared modality adapter and the third general adapter. The output result of the third visible light modality adapter is superimposed on the first output end of the third general adapter. The third visible light modality adapter is connected to the third thermal infrared modality adapter through a meta-learner. The output result of the third thermal infrared modality adapter is superimposed on the second output end of the third general adapter through a dimensionality reduction unit. When only the output result of the first output end of the third general adapter is available, this output result is output through the second instance adapter. Otherwise, the results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
[0039] Continue to refer to Figure 1, the first visible light modality adapter and the first thermal infrared modality adapter have the same structure, both consisting of a 3×3 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the second visible light modality adapter and the second thermal infrared modality adapter have the same structure, both consisting of a 1×1 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the third visible light modality adapter, the third thermal infrared modality adapter, and the dimensionality reduction unit have the same structure, both consisting of a 1×1 convolutional layer, ReLU, and LRN cascaded in sequence; the first general adapter consists of a 7×7 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the second general adapter consists of a 5×5 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the third general adapter consists of a 3×3 convolutional layer, ReLU, and LRN; the meta-learner consists of two learning units cascaded in sequence, both learning units consisting of ReLU and a fully connected layer, and the dimensionalities of the fully connected layers of the two learning units are different; the first instance adapter consists of a fully connected layer FC4, a ReLU, a dropout function, a fully connected layer FC5, another ReLU, another dropout function, and a fully connected layer FC6 cascaded in sequence, and the second instance adapter has the same structure as the first instance adapter, but the dimensionality of the fully connected layer FC4 is different.
[0040] The training and tracking processes of the multi-modal visual tracking control model of the present invention are introduced in detail below. The training process includes:
[0041] Use the parameters of the first three layers of the MDNet network pre-trained on the Imagene-Vid dataset and the domain-shared fully connected layer as the pre-trained parameters in the network, and the parameters used for each convolutional kernel group of each layer are the same. Then, initialize the parameters of the domain-specific layer, modality adapter, and meta-learner using a Gaussian distribution. And intercept the first three convolutional layers as the backbone network, and copy the channel parameters of the last three fully connected layers as the initialization parameters of the general adapter, where the sizes of the three convolutional layers are 7*7*96, 5*5*256, and 3*3*512 respectively. Each convolutional layer is followed by an activation function ReLu, and the first two layers also have a local response function LRN and a max pooling layer function MaxPool. The instance adapter is a fully connected layer that converts the dimension 512-216-2 to 1024-512-2.
[0042] Train the entire network model using the calibrated visible light and infrared dataset (visible light images and their corresponding thermal infrared images) and the visible light-only dataset. Assume there are K labeled sequences in the dataset, then establish K fc6 ilayers, where fc6 is a fully connected layer, and each fully connected layer corresponds to a video sequence. For different datasets, the instance configurator used for the visible light and infrared dataset is not the same as that used for the visible light only dataset, because their channel dimensions are different, and the challenges and target types in different datasets are also different. Using different instance configurators helps the network learn a better general adapter and visible light modality adapter. Moreover, the ways for the visible light and infrared dataset and the visible light only dataset to pass through the modality adapter are different. The network structure that the visible light and infrared dataset passes through is the general adapter, the visible light modality adapter, the thermal infrared modality adapter, the meta-learner, and the specific instance adapter. The network structure that the visible light only modality dataset passes through is the general adapter, the visible light modality adapter, and its specific instance adapter. The input order of the two datasets into the network is also different. First, use the visible light and infrared modality dataset to iterate about 70 times to pre-train the network model, and then alternately input the visible light and infrared dataset and the visible light only dataset into the network. At this time, the meta-learner starts to learn when the visible light and infrared dataset is input.
[0043] Select S according to the given ground truth box in each frame + = 4 (IOU≥0.7) and S - = 12 (IOU≤0.5) of the number of samples. Where S + represents positive samples, S - represents negative samples, and IOU represents the intersection over union between the collected samples and the ground truth box. Through the collected positive and negative samples, use the visible light and infrared modality dataset to iterate 70 times to pre-train the network model. Each iteration is processed according to the following method: The minibatch (batchsiz = 128) in the k-th iteration is randomly selected from 8 frames of visible light and infrared images in the k-th video sequence, which contains 32 positive samples and 96 negative samples, and activate the corresponding last fully connected layer. And set the learning rate of the convolutional layer to 0.0001 and the learning rate of the fully connected layer to 0.0002. Then randomly cross-input the visible light and infrared dataset and the visible light only modality dataset into the network for 130 iterations, and the learning rate and sample sampling method are the same as the above method. Save the finally obtained trained model for the target tracking process.
[0044] As Figure 2 shown, the tracking process includes:
[0045] During the tracking process, the dataset of only visible light modality and the meta-learner will no longer be used. First, a new fc6 layer is constructed. The learning rate of the convolutional layer is fixed, the learning rate of the fc6 layer is set to 0.001, and the learning rate of the remaining fully connected layers is set to 0.0005. A total of thirty iterations are performed. From the first frame of paired multi-modal images provided by the tracking video sequence and the ground truth box framing the target area, 5500 samples S are randomly generated according to the following strategy + = 500 (IOU ≥ 0.7) and S - = 5000 (IOU ≤ 0.3). These samples are used to initialize the tracking model and perform bounding box regression training, that is, the sample boxes are translated and scaled so that the values after linear regression are very close to the true values. That is, from the first frame of the given test sequence, a simple linear regression model is trained using the features of the output of the last convolutional layer of the samples near the target position to predict the exact position of the target. This bounding box regression training is only performed in the first frame. During training in subsequent frames, if the sample to be evaluated is reliable f + (x i ) ≥ 0.5, the regression model is used to evaluate the equation to adjust the position of the target. The specific parameter setting is S + = 1000 (IOU ≥ 0.6). These samples are set to a minibatch with a batchsize of 256 and trained for a total of five iterations, and the boundary regression weight parameters are fine-tuned. In addition, the samples in the first frame are added to the positive and negative sample sets.
[0046] After the initialization training and bounding box regression are completed. Using the target position of the previous frame as the mean and (0.09r 2 , 0.09r 2 , 0.25) as the covariance, 256 pairs of candidate samples are generated, where: r is the average of the width and height of the target box in the previous frame.
[0047] These samples are respectively fed into the corresponding adapters. For the visible light samples, they are input into the visible light modality adapter and the general modality adapter. The visible light modality adapter consists of a group of small-sized convolutional layers, namely three convolutions parallel to each convolutional layer of the general adapter: 3*3*96, 1*1*256, and 1*1*512. After each convolution, there is a ReLU function and an LRN function, and at the end of the first two convolutional layers, there is a max pooling function MaxPool with a size of 5*5. The thermal infrared samples are input into the thermal infrared modality adapter and the general adapter, and the structure of the thermal infrared modality adapter is exactly the same as that of the visible light modality adapter described above. Then, the output features of the two modalities in the general adapter are respectively matrix-added and fused with the output features of the modality adapters of their respective modalities, and then fed into the next-level adapter. Finally, the two modality feature maps obtained are concatenated based on the channel dimension using the concatenate function to obtain a final feature map, that is, the original two 512*3*3 are concatenated into a 1024*3*3, and then input into the instance adapter. Its structure is composed of the sum of three Dropout (random inactivation units) functions and three fully connected layers, with respective dimensions of 1024*512, 512*512, and 512*2, and at the end of the first two fully connected layers, there is an activation function ReLU to form the entire instance adapter. Through the last layer with a softmax layer, the scores of each candidate sample being judged as a positive sample and a negative sample are obtained, denoted as f + (x i ) and f - (x i ), and the target position of the next frame is generated by the following formula: where x i represents the i-th sampled sample, f + (x i ) represents the obtained positive sample score, and f - (x i ) represents the obtained negative sample score. x * is the predicted target position.
[0048] According to the candidate box using the formula, calculate the position of the target to be recognized in the current frame pair. Among them, x * is the position of the target to be recognized in the current frame pair; is the function to find the variable at the maximum value; f + (x i ) is the positive sample score in the current frame pair; select the optimal candidate box x * ;
[0049] If the probability of the optimal candidate box obtained in the previous frame is greater than 0.5, it is determined that the tracking is successful. Positive and negative samples are obtained according to the IOU around the obtained optimal candidate box, 50 positive samples (IOU≥0.6) and 200 negative samples (IOU≤0.3). And these samples are saved as the convolved features. Among them, the total positive sample set saves the positive samples of the most recent 100 tracking successful frames, and the total negative sample set saves the negative samples of the most recent 20 tracking successful frames.
[0050] If the amount of data in the short-term set is greater than 20, the frame with the smallest number of frames is deleted. If the number of the long-term set is greater than 100, the frame with the smallest number of frames is deleted. And the optimal candidate box is adjusted by bounding box regression. If the probability of the optimal candidate box obtained in the previous frame is less than 0.5, it is determined that the tracking fails. Positive and negative samples collected in the most recent 20 frames of successful tracking and positive and negative samples generated in the first frame are extracted from the positive and negative sample sets to form a training set. If the current frame is a multiple of 10, the data in the long-term updated set is used to optimize and update the last layer of the network.
[0051] The present invention conducts simulation experiments, tests on the publicly available datasets GTOT and RGBT234 respectively, and evaluates the test results with other trackers in terms of SR (success rate) and PR (precision). The experimental results are as Figures 3 to 6 shown, Figure 3 and Figure 4 are the test results on the dataset GTOT, Figure 5 and Figure 6 are the test results on RGBT234. The success rate of the present invention tested on the publicly available dataset GTOT is very high compared with the success rates of other trackers, reaching 0.715, and the precision is the highest compared with other trackers, reaching 0.895. The success rate of the present invention tested on the publicly available dataset RGBT234 is the highest compared with the success rates of other trackers, and the precision is also the highest compared with other trackers, reaching 0.833. Therefore, the multi-modal visual tracking control model and its tracking method provided by the present invention have a very high success rate and precision, and the performance is significantly improved compared with the existing trackers.
[0052] Through the above technical solutions, in the process of training the multi-modal visual tracking control model of the present invention, image pairs are first input into the model for iterative training for the first preset number of times, and then separate visible light images and image pairs are alternately input into the model to continue training for the second preset number of times. The paired image dataset is relatively small, and additional visible light images are added to participate in the training to make up for the deficiency of the training set. At the same time, a meta-learner is introduced to further solve the problem of insufficient network training caused by insufficient datasets in the prior art, avoiding the model from converging too quickly or not converging due to a small dataset, thereby avoiding inaccurate training results and improving the model performance, and solving the problem that the prior art limits the improvement of the model performance due to limited data volume.
[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A visible light and infrared visual tracking method based on meta - learning parameter transfer, characterized in that, The method includes: Step 1: Construct a multi-modal visual tracking control model; Step 2: Input the samples into the multi-modal visual tracking control model for training. The samples include multiple image pairs formed by multiple thermal infrared images and their corresponding visible light images, as well as multiple individual visible light images. During the training process, first input the image pairs into the model for the first preset number of iterative trainings, and then alternately input the individual visible light images and the image pairs to continue training the model for the second preset number of times. And when using the image pairs for model training during the alternating training process, the images need to pass through the meta-learner; Step 3: After the training is completed, collect the video in real time, extract the paired thermal infrared images and their corresponding visible light images, input them into the trained model, and the model tracks and outputs the predicted target position; The multi-modal visual tracking control model includes the first general adapter to the third general adapter, the first visible light modal adapter to the third visible light modal adapter, the first thermal infrared modal adapter to the third thermal infrared modal adapter, the meta-learner, the first instance adapter and the second instance adapter numbered in sequence. The first visible light modal adapter is connected to an input end of the first general adapter and receives the individual visible light image or the visible light image corresponding to the thermal infrared image. The first thermal infrared modal adapter is connected to the other input end of the first general adapter and receives the thermal infrared image. The output result of the first visible light modal adapter is superimposed on the first output end of the first general adapter and then input into the second visible light modal adapter and the second general adapter. The output result of the first thermal infrared modal adapter is superimposed on the second output end of the first general adapter and then input into the second thermal infrared modal adapter and the second general adapter. The output result of the second visible light modal adapter is superimposed on the first output end of the second general adapter and then input into the third visible light modal adapter and the third general adapter. The output result of the second thermal infrared modal adapter is superimposed on the second output end of the second general adapter and then input into the third thermal infrared modal adapter and the third general adapter. The output result of the third visible light modal adapter is superimposed on the first output end of the third general adapter. The third visible light modal adapter is connected to the third thermal infrared modal adapter through the meta-learner. The output result of the third thermal infrared modal adapter is superimposed on the second output end of the third general adapter through the dimensionality reduction unit. When only the output result of the first output end of the third general adapter is output, this output result is output through the second instance adapter. Otherwise, the results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter; The first visible light modality adapter and the first thermal infrared modality adapter have the same structure, both consisting of a 3×3 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the second visible light modality adapter and the second thermal infrared modality adapter have the same structure, both consisting of a 1×1 convolutional layer, ReLU, LRN, and a 5×5 max pooling layer cascaded in sequence; the third visible light modality adapter, the third thermal infrared modality adapter, and the dimensionality reduction unit have the same structure, both consisting of a 1×1 convolutional layer, ReLU, and LRN cascaded in sequence; the first general adapter consists of a 7×7 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the second general adapter consists of a 5×5 convolutional layer, ReLU, LRN, and a 3×3 max pooling layer; the third general adapter consists of a 3×3 convolutional layer, ReLU, and LRN; the meta-learner consists of two learning units cascaded in sequence, and both learning units consist of ReLU and a fully connected layer, and the dimensions of the fully connected layers of the two learning units are different; the first instance adapter consists of a fully connected layer FC4, a ReLU, a dropout function, a fully connected layer FC5, another ReLU, another dropout function, and a fully connected layer FC6 cascaded in sequence, and the structure of the second instance adapter is the same as that of the first instance adapter, but the dimension of the fully connected layer FC4 is different; The training process of the multi-modal visual tracking control model includes: Iteratively pre-training the network model 70 times using multiple image pairs formed by thermal infrared images and their corresponding visible light images, setting the batch size to 128, the learning rate of the convolutional layer to 0.0001, and the learning rate of the fully connected layer to 0.0002. During the training process, the visible light image and its corresponding thermal infrared image are input into the model simultaneously. The visible light image passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position in sequence, and then the output result is fused to the first output end of the third general adapter. The thermal infrared image passes through the first thermal infrared adapter and the first general adapter at the corresponding position, the second thermal infrared adapter and the second general adapter at the corresponding position, and the third thermal infrared adapter and the third general adapter at the corresponding position in sequence, and then the output result is fused to the second output end of the third general adapter. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
2. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 1, characterized in that, The sample acquisition method used during training is as follows: In each frame of the video, select S according to the given ground truth box + = 4, the number of samples with IOU ≥ 0.7 and S - = 12, the number of samples with IOU ≤ 0.5, where S + represents positive samples, S - represents negative samples, IOU represents the intersection over union between the collected samples and the ground truth box, through the collected positive and negative samples.
3. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 2, characterized in that, The training process of the multi-modal visual tracking control model also includes: alternately inputting single visible light images and image pairs to continue training the model 130 times, with the learning rate and sample sampling method being the same as those for iteratively pre-training the network model 70 times using thermal infrared images and their corresponding visible light images.
4. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 3, characterized in that, Each time when training the model with a single visible light image, the single visible light image sequentially passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position, and then the output result is fused to the first output end of the third general adapter, and the output result of the first output end of the third general adapter is output from the second instance adapter.
5. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 3, characterized in that, During the process of alternately inputting a single visible light image and an image pair to continue training the model 130 times, each time when training the model with the image pair, the visible light image sequentially passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position, and then the output result is fused to the first output end of the third general adapter; the thermal infrared image sequentially passes through the first thermal infrared adapter and the first general adapter at the corresponding position, the second thermal infrared adapter and the second general adapter at the corresponding position, and the output results of the second thermal infrared adapter and the second general adapter at the corresponding position are fused and then input to the third thermal infrared adapter and the third general adapter. The output of the third visible light adapter is input to the third thermal infrared adapter through the meta-learner. The output result of the third thermal infrared adapter is fused to the second output end of the third general adapter after passing through the dimensionality reduction unit. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter.
6. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 1, characterized in that, Step 3 includes: after the training is completed, collecting videos in real time, extracting paired thermal infrared images and their corresponding visible light images and inputting them into the trained model. The visible light image sequentially passes through the first visible light adapter and the first general adapter at the corresponding position, the second visible light adapter and the second general adapter at the corresponding position, and the third visible light adapter and the third general adapter at the corresponding position, and then the output result is fused to the first output end of the third general adapter; the thermal infrared image sequentially passes through the first thermal infrared adapter and the first general adapter at the corresponding position, the second thermal infrared adapter and the second general adapter at the corresponding position, and the third thermal infrared adapter and the third general adapter at the corresponding position, and then the output result is fused to the second output end of the third general adapter. The results of the first output end and the second output end of the third general adapter are fused and then output through the first instance adapter to output the predicted target position.
7. The visible light and infrared visual tracking method based on meta - learning parameter transfer according to claim 6, characterized in that, The output through the first instance adapter to output the predicted target position includes: The fully connected layer FC6 of the first instance adapter contains a softmax layer that calculates the positive and negative scores for each sample feature: f + (x i ) and f - (x i ). The predicted target position is obtained through the formula . Here, x i represents the i-th sampled sample, f + (x i ) represents the obtained positive sample score, f - (x i ) represents the obtained negative sample score, and x * is the predicted target position.
Citation Information
Patent Citations
RGBT target tracking method based on cross-modal sharing and specific representation form
CN113077491A
RGBT visual tracking method and system based on two-stage fusion structure search
CN113837296A