A Dual-Modal Visual Tracking Method Based on High-Rank Features and Position Attention

By introducing target position attention module and high-rank guidance module into the backbone network of the RGBT tracking method, the problem of poor tracking performance in the prior art under harsh environments is solved, and more effective fusion of visible and thermal infrared features and noise reduction are achieved.

CN114022516BActive Publication Date: 2025-05-27ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111346472.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2025-05-27
Estimated Expiration
2041-11-15

AI Technical Summary

Technical Problem

Existing RGBT tracking methods perform poorly in harsh environments and are difficult to effectively fuse visible and thermal infrared features, resulting in increased computational burden and noise introduction.

Method used

A dual-modal visual tracking method based on high-rank features and position attention is adopted. By introducing a target position attention module and a high-rank guidance module in the backbone network, the target position information is paid attention to the target position information and guide the fusion of visible light and thermal infrared feature maps.

Benefits of technology

Improve the effect of target tracking, reduce the impact of noise, and optimize the computational burden to achieve more stable signal fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022516B_ABST
    Figure CN114022516B_ABST
Patent Text Reader

Abstract

The present invention discloses a dual-modal visual tracking method based on high-rank features and position attention, and provides a dual-modal visual tracking method based on high-rank features and position attention. By introducing a target position attention module into the backbone network to focus on target position information, and using a high-rank guidance module to focus on important channels and guide the fusion of visible light and thermal infrared feature maps, the effect of target tracking is further improved, and the network model can be judged whether to be updated according to the success or failure of the target result. The present invention can more accurately locate the position of the target while reducing noise interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to computer vision technology, and in particular relates to a dual-modal visual tracking method based on high-rank features and position attention. Background Art

[0002] Visual target tracking is an important task in computer vision and has wide applications in many fields, such as military, intelligent transportation and video surveillance.

[0003] In recent years, great progress has been made in visual object tracking, especially RGB (visible light) tracking. However, the performance of RGB tracking is not satisfactory under harsh environmental conditions, such as low illumination, rain, and smoke. Thermal infrared sensors provide more stable signals for these situations, and thermal cameras have recently become more affordable, so thermal infrared images have been applied to many computer vision tasks. Thermal sensors are based on thermal radiation from blood vessels on the surface of the human body or the heat distribution of any object. They are insensitive to changes in illumination, camouflage and posture changes of objects, and have a strong ability to penetrate smoke and haze, so they can provide strong complementary information to visible cameras. However, compared with visual sensors, thermal sensors have high image noise and low resolution, and poor edge and texture information. Therefore, RGBT (visible light and thermal infrared) tracking has received widespread attention in recent years and has made great progress.

[0004] However, there are still many problems to be solved in effectively integrating RGB and thermal infrared, such as how to integrate RGB and thermal infrared to make full use of their complementarity.

[0005] At present, there are two main aspects of RGBT tracking methods. On the one hand, how to design a suitable representation learning framework for RGBT tracking, the existing algorithm proposes a cross-modal sorting algorithm to calculate the importance weight of each patch, and then constructs a robust RGBT feature description of the target object; on the other hand, how to achieve adaptive fusion of different modes for RGBT tracking, the existing algorithm is based on the collaborative sparse representation under the Bayesian filtering framework, by optimizing the reliability weight of each modality for online fusion, or according to the classification score, using the maximum threshold principle to optimize the modality weight.

[0006] The above methods still have some shortcomings, as follows:

[0007] (1) The existing technology is to adaptively fuse the features of visible light and thermal infrared in an end-to-end manner. Usually, the visible light and thermal infrared modalities are modeled separately first, and then the weight of each modality is adaptively learned to fuse the two modalities. The adaptive online calculation method channel will increase the computational burden of the algorithm model and ignore the characteristics of the filter itself.

[0008] (2) Existing technologies enhance feature extraction capabilities by introducing shared features and unique features, but they do not pay attention to the large amount of redundant information between features, which often easily introduces noise. Summary of the invention

[0009] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and to provide a dual-modal visual tracking method based on high-rank features and position attention. The method introduces a target position attention module into the backbone network to focus on the target position information, and uses a high-rank guidance module to focus on important channels and guide the fusion of visible light and thermal infrared feature maps, thereby further improving the target tracking effect.

[0010] Technical solution: A dual-modal visual tracking method based on high-rank features and position attention of the present invention comprises the following steps:

[0011] Step 1: For the registered multimodal image, take the first frame of the video in its corresponding visible light and thermal infrared video, manually frame the target frame to be tracked on the first frame, and then perform Gaussian distribution sampling with the center point of the target frame as the mean, and collect a total of several (for example, 256) candidate sample frames;

[0012] Step 2: Input the candidate sample frames of the two modalities obtained in step 1 into the network model respectively, and extract features from the candidate sample frames of the two modalities through the backbone network of the network model.

[0013] The backbone network uses the first three convolutional layers of VGG-M. A branch is added to each of these three convolutional layers, through which the target position attention module is introduced to focus on the position information of the tracking target.

[0014] Among them, for the first convolutional layer, the feature maps of the visible light and thermal infrared modes are directly added and then sent to the target position attention module;

[0015] For the second convolutional layer, convolution and pooling operations are introduced on the branch of the target position attention module. Through convolution and pooling, the branch feature map of the target position attention module here is made to match the feature map size of the backbone network;

[0016] Step 3: After the third convolution operation, a high-rank guidance module is introduced into the backbone network of both visible light and thermal infrared modalities to guide the fusion of the two modalities, and the feature map corresponding to the noise channel is deleted;

[0017] Step 4: Send the fused features (feature maps after cat) of the high-rank guidance module to three fully connected layers. There are three fully connected layers in total. The first two fully connected layers are followed by a neuron random activation function to alleviate the overfitting problem. The third fully connected layer is used to classify whether the sample frame is a positive sample or a negative sample, and a softmax layer is introduced after the third fully connected layer. The positive and negative sample scores of the candidate sample frames are obtained through softmax calculation. The candidate frame with the highest score in the positive sample is predicted as the target result to be tracked.

[0018] Step 5: Determine whether to update the model based on the success or failure of the target tracking result obtained above. If the tracking fails, perform a short-term update; if the tracking is successful, continue to track the next frame of the image; and perform a long-term update every ten frames of images.

[0019] Furthermore, the backbone network in step 2 uses the first three convolutional layers of VGG-M, and the convolution kernel sizes of these three convolutional layers are 7x7, 5x5, and 3x3, respectively.

[0020] Furthermore, the specific process of the high-rank guidance module in step 3 guiding the fusion of the two modal features of visible light and thermal infrared is as follows:

[0021] First, the rank information corresponding to the feature maps obtained by the third convolutional layer of the two modal images is calculated respectively, and then the ranks of the two modalities are normalized respectively, and the feature maps with rank values ​​lower than the set threshold are set to zero. Then, the two normalized rank values ​​are used as weights to guide the feature fusion of the visible light and thermal infrared modalities;

[0022] Here, the feature fusion method is to perform a cat operation on the feature maps of visible light and thermal infrared images, that is, link them according to the first dimension (up and down). Here, high-rank information is used for selection to delete redundant information, so as to reduce the impact of noise on the network.

[0023] Furthermore, the numbers of channels of the three fully connected layers in step 4 are 1024, 512 and 2 respectively.

[0024] Furthermore, the neuron random activation function in step 4 adopts the Dropout function, which is selected as a trick for training deep neural networks. In each training batch, half of the feature detectors are ignored (half of the hidden layer node values ​​are set to 0), thereby significantly reducing overfitting.

[0025] Furthermore, in step 5, when the obtained target result score is greater than zero, the tracking is considered to be successful; when the obtained target result score is less than zero, the tracking is considered to have failed.

[0026] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0027] (1) The present invention uses VGG-M as the backbone network to extract features, and introduces a target position attention module in the convolutional layer. The target position attention module is used to focus on the position information of the tracked target, which is more conducive to locating the target.

[0028] (2) The present invention also introduces a high-rank feature guidance module after the third convolutional layer. The high-rank feature guidance module ranks the importance of channels, that is, it pays attention to the different importance of different channels in different modalities, which is more conducive to the fusion of visible light and thermal infrared images.

[0029] (3) In the present invention, a zeroing operation is used for feature maps with smaller rank to alleviate the noise problem caused by low-quality feature maps, thereby achieving better target tracking effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0031] Figure 2 A schematic diagram of a network model structure in an embodiment;

[0032] Figure 3 is a visible light image in the embodiment;

[0033] Figure 4 is the corresponding thermal infrared image in the embodiment;

[0034] Figure 5 Schematic diagram of rank information obtained in the embodiment;

[0035] Figure 6 A schematic diagram of the comparison of the embodiment on the data set GTOT;

[0036] Figure 7 A schematic diagram of comparison of the embodiment on the data set RGBT234;

[0037] in, Figure 6 (a) and Figure 7 (a) is the position error threshold curve diagram, Figure 6 (b) and Figure 7 (b) is the overlap threshold comparison chart. DETAILED DESCRIPTION

[0038] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0039] Embodiment 1:

[0040] like Figure 1As shown, a dual-modal visual tracking method based on high-rank features and position attention in this embodiment includes the following steps:

[0041] Step 1: For the registered multimodal image, take the first frame of the video in its corresponding visible light and thermal infrared video respectively, manually frame the target frame to be tracked on the first frame, and then perform Gaussian distribution sampling with the center point of the target frame as the mean, and collect a total of several candidate sample frames;

[0042] Step 2: Input the candidate sample frames of the two modalities obtained in step 1 into the network model respectively, such as Figure 2 As shown in the figure, the feature extraction of the candidate sample frames of the two modalities is performed through the backbone network of the network model. The backbone network uses the first three convolutional layers of VGG-M. A branch is added to each of the three convolutional layers. The target position attention module is introduced through the branch to focus on the position information of the tracking target.

[0043] Among them, for the first convolutional layer, the feature maps of the visible light and thermal infrared modes are directly added and then sent to the target position attention module;

[0044] For the second convolutional layer, convolution and pooling operations are introduced on the branch of the target position attention module. Through convolution and pooling, the branch feature map of the target position attention module here is made to match the feature map size of the backbone network;

[0045] Step 3: After the third convolution operation, a high-rank guidance module is introduced into the backbone network of both visible light and thermal infrared modalities to guide the fusion of the two modalities, and the feature map corresponding to the noise channel is deleted;

[0046] Step 4: Send the fused features (feature maps after cat) of the high-rank guidance module to three fully connected layers. There are three fully connected layers in total. The first two fully connected layers are followed by a neuron random activation function to alleviate the overfitting problem. The third fully connected layer is used to classify whether the sample frame is a positive sample or a negative sample, and a softmax layer is introduced after the third fully connected layer. The positive and negative sample scores of the candidate sample frames are obtained through softmax calculation. The candidate frame with the highest score in the positive sample is predicted as the target result to be tracked.

[0047] Step 5: Determine whether to update the model based on the success or failure of the target tracking result obtained above. If the tracking fails, perform a short-term update; if the tracking is successful, continue to track the next frame of the image; and perform a long-term update every ten frames of images.

[0048] Embodiment 2:

[0049] Here are two modal diagrams such as 3 and Figure 4This embodiment includes two processes, training and testing, and the specific steps are as follows:

[0050] (1) Network training process:

[0051] (1.1) Since the rank needs to be calculated offline, the network is trained twice here. First, the pre-trained model is loaded for the parameters of the first three convolutional layers of VCC-M. The backbone network of this embodiment has three convolutional layers, and the convolution kernel sizes of each convolutional layer are 7x7, 5x5 and 3x3 respectively. Each convolutional layer contains an activation function Relu layer. The first convolutional layer and the second convolutional layer also contain a local response function LRN layer and a maximum pooling function layer. There is a branch on each of the three convolutional layers of this embodiment, which is used to introduce the target position attention module. The convolution size of each branch is set to match the size of the backbone feature map.

[0052] (1.2) During the first training, a true value box is manually annotated on each frame of the image, and the manually annotated true value box is used to train the entire network. During the specific training, 256 candidate sample boxes are selected near the true value box. These 256 candidate sample boxes are divided into positive samples and negative samples according to the intersection-over-union (IOU) ratio between the true value box and the candidate sample. When the IOU is greater than or equal to 0.7, it is considered a positive sample, and when the IOU is less than or equal to 0.5, it is considered a negative sample.

[0053] (1.3) The network is trained using stochastic gradient descent (SGD) and cross entropy loss for 100 epoch iterations. In each iteration, 8 frames are randomly selected from each video sequence, and then 64 positive samples and 192 positive samples are selected from each frame. For the judgment of positive samples, samples with an IOU greater than 0.7 with the true value frame are judged as positive samples, and samples with an IOU less than 0.5 with the true value frame are judged as negative samples. Multi-domain training is performed during training.

[0054] (1.4) Second training: After the network model is obtained in the first training, 5 video sequences are randomly selected here, and the rank of the feature map of the second picture in these 5 sequences is tracked and calculated. The reason why the first picture is not selected is that the algorithm performs difficult negative sample mining on the first picture. The information of the feature map rank of the second picture of the 5 video sequences is saved here, and then its average is calculated. Then the average rank of the feature map is multiplied to the feature map as a weight, and then the network is trained again. The setting of the hyperparameters is basically the same as the first training, the only difference is the number of iterations, which is 500 for the second training.

[0055] In the above steps (1.2) and (1.3), other numbers of candidate sample frames may be selected. This embodiment uses 256 candidate frames based on the Manet algorithm. At the same time, the ratio of positive and negative samples can be 1:3.

[0056] (2) Network tracking process:

[0057] (2.1) In the tracking video, the true value frame of the target to be tracked is given in the first frame, and then 500 positive samples and 5000 negative samples are sampled. 30 iterations of training are performed when tracking the first frame. Then these 5500 positive and negative samples are used to train the network model to obtain a new fc6 layer. At this time, the learning rate of the convolution layer is fixed, the learning rate of the first and second fully connected layers is set to 0.0005, and the learning rate of the last fully connected layer is set to 0.001. After the initialization work is completed, the position of the target in the previous frame is averaged, and then Gaussian distribution sampling is used to take 256 candidate sample frames.

[0058] (2.2) The candidate sample frame is sent to the backbone network, and a branch is added to each convolution layer to introduce the target position attention module, so that the network can better locate the target position. The high-rank feature guidance fusion module is not used during tracking, because the average rank of the feature map is calculated offline, and the information of the saved feature map rank is used as the weight to guide the modal fusion; the feature map after rank guidance fusion is sent to the fully connected layer. The fully connected layer has three layers, and a softmax layer is connected after the last fully connected layer to obtain the scores of positive and negative samples.

[0059] (2.3) When the target result score predicted by the network model is greater than zero, tracking is considered successful; when the target result score predicted by the model is less than zero, tracking is considered unsuccessful. When tracking is successful, positive and negative samples are collected in the current frame, mainly 50 positive samples and 200 negative samples, and these 250 sample frames are added to the positive and negative sample sets. When the number of frames in the positive and negative sample sets is greater than 100, the positive sample frames of the earliest frame are discarded. If the number of frames is greater than 20, the negative sample frames of the earliest frame are discarded.

[0060] When tracking fails, a short-term update of the network model is required: 32 positive sample frames and 96 negative sample frames are extracted from the positive and negative sample sets to fine-tune the parameters of the fully connected layer.

[0061] (2.4) When the network model performs online tracking, it will not only perform short-term updates when the previous tracking fails, but will also automatically perform long-term updates every 10 frames. The long-term update method is the same as the short-term update method. If the network model does not meet the requirements of either long-term or short-term updates, it will directly perform target tracking for the next frame.

[0062] As shown in Table 1 and Table 2, in this embodiment, the accuracy and success rate of the technical solution of the present invention are compared with other prior arts.

[0063] Table 1 Results on the GTOT dataset

[0064]

[0065] Table 2 Results on the RGBT234 dataset

[0066]

[0067] Here, accuracy is the percentage of frames where the distance between the output location box and the ground-truth bounding box is below a predefined threshold; success rate is the percentage of frames where the overlap between the output bounding box and the ground-truth bounding box is greater than the threshold.

[0068] like Figure 6 and Figure 7 As shown, this embodiment uses different linearity to describe a schematic diagram comparing the position error threshold and overlap threshold between the technical solution of the present invention and the prior art solution on different data sets. Figure 6 and Figure 7 The four figures show that the technical solution of the present invention is superior to all the currently published RGBT tracking algorithms in terms of accuracy.

Claims

1. A dual-modal visual tracking method based on high-rank features and position attention, characterized in that: It includes the following steps: Step 1: For the registered multi-modal images, take the first frame images of the corresponding visible light and thermal infrared videos respectively. Frame the target box on the first frame, and then perform Gaussian distribution sampling with the center point of the target box as the mean, and collect a number of candidate sample boxes in total; Step 2: Input the candidate sample boxes of the two modalities obtained in Step 1 into the network model respectively, and extract features of the candidate sample boxes of the two modalities through the backbone network of the network model. The backbone network uses the first three convolutional layers of VGG-M, and a branch is added to each of these three convolutional layers. Through this branch, a target position attention module is introduced to focus on the position information of the tracking target; Among them, for the first convolutional layer, directly perform an addition operation on the feature maps of the visible light and thermal infrared modalities, and then send them into the target position attention module; For the second convolutional layer, convolution and pooling operations are introduced on the branch of its target position attention module. Through convolution and pooling, the size of the branch feature map of the target position attention module here is matched with the feature map of the backbone network; Step 3: After the third convolutional operation, a high-rank guidance module is introduced for the backbone networks of both the visible light and thermal infrared modalities. The high-rank guidance module guides the fusion of the two modalities and deletes the feature maps corresponding to the noise channels at the same time. The specific process is as follows: First, calculate the rank information corresponding to the feature maps obtained by the third convolutional layer of the two-modal images respectively, and then perform a normalization operation on the ranks of the two modalities. For the feature maps with rank values lower than the set threshold, perform a zeroing operation. Then, use the two normalized rank values as weights to guide the feature fusion of the visible light and thermal infrared modalities. Here, the feature fusion method is to perform a concatenation (concat) operation on the feature maps of the visible light and thermal infrared images; Step 4: Send the feature fusion passed through the high-rank guidance module into three fully connected layers. There are a total of three fully connected layers. Neuron random activation functions are added after the first two fully connected layers to alleviate the problem of overfitting. The third fully connected layer is used to distinguish whether the sample box is a positive sample or a negative sample, and a softmax layer is introduced after the third fully connected layer. After softmax calculation, the positive and negative sample scores of the candidate sample boxes are obtained. The candidate box with the highest score among the positive samples is predicted as the target result to be tracked; Step 5: Judge whether to update the network model according to the success or failure of the target result to be tracked obtained above. If the tracking fails, perform a short-term update; if the tracking is successful, continue to track the next frame of the picture; and perform a long-term update every ten frames of images.

2. The dual-modal visual tracking method based on high-rank features and position attention according to claim 1, characterized in that: In Step 2, the backbone network uses the first three convolutional layers of VGG-M, and the convolutional kernel sizes of these three convolutional layers are 7x7, 5x5, and 3x3 respectively.

3. The dual-modal visual tracking method based on high-rank features and position attention according to claim 1, characterized in that: The number of channels of the three fully connected layers in step 4 is 1024, 512, and 2 respectively.

4. The dual-modal visual tracking method based on high-rank features and position attention according to claim 1, characterized in that: In step 4, the neuron random activation function adopts the Dropout function, which is used as a trick for training the deep neural network for selection. In each training batch, overfitting is reduced by ignoring half of the feature detectors.

5. The dual-modal visual tracking method based on high-rank features and position attention according to claim 1, characterized in that: In step 5, when the obtained target result score is greater than zero, the tracking is considered successful; when the obtained target result score is less than zero, the tracking is considered failed.

Citation Information

Patent Citations

  • Training and visible light infrared visual tracking method based on adapter mutual learning model

    CN110874590A

  • Multi-source image fusion method based on low-rank decomposition and convolution sparse coding

    CN111833284A