A Motion-Aware Self-Supervised RGBT Tracking Method Based on Multimodal Hierarchical Transformer

Through the motion-aware self-supervised RGBT tracking method of multimodal hierarchical Transformer, the MHTF module and the MAM module capture the long-distance dependence of visible light and thermal infrared images is solved, and challenges such as occlusion and rapid motion are achieved, more precise feature fusion and classification are achieved, and the effect of self-supervised RGBT tracking is improved.

CN117197187BActive Publication Date: 2025-07-18CHINA UNIV OF MINING & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311169927.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-12
Publication Date
2025-07-18
Estimated Expiration
2043-09-12

AI Technical Summary

Technical Problem

The existing RGBT tracking methods have challenges in occlusion, fast target motion and large-scale transformation. The self-supervised tracker has insufficient feature robustness due to the lack of accurate manual annotation. The existing Transformer fusion method has a large amount of computation and does not fully utilize the complementary information of visible light and thermal infrared images.

Method used

The motion-aware self-supervised RGBT tracking method of multimodal hierarchical Transformer is adopted to capture the long-distance dependence between visible light and thermal infrared images through the MHTF module, and the motion vector is recorded in combination with the MAM module, and the multi-head cross-attention mechanism is used to perform feature fusion and classification. The self-supervised training method is used to reduce the need for manual annotation.

Benefits of technology

Without increasing the amount of calculation, the image complementary information is fully utilized to achieve more accurate feature fusion and classification, which improves the accuracy of target tracking, reduces the dependence on manual annotation, and helps the development of self-supervised RGBT tracking technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197187B_ABST
    Figure CN117197187B_ABST
Patent Text Reader

Abstract

The present invention discloses a motion-aware self-supervised RGBT tracking method based on a multi-modal hierarchical Transformer. First, ResNet50 is used to extract the features of RGB images and thermal infrared images. Then, the MHTF module is used to capture the long-distance dependencies between the two modal features on the channel, and a convolution-based cross-correlation operation is performed on the fused features. The classification-enhanced score map based on the multi-head cross-attention mechanism is used to assist in achieving accurate classification. The MAM module is introduced to record the search frame features and extract the corresponding motion vectors, and these vectors are used during the training of the network model to strengthen the consistency with the current search frame features, minimizing the losses obtained from the cross-correlation operation and the MAM module. Finally, the video frames are input into the trained network model for tracking to obtain the tracking results. The method of the present invention makes full use of the complementary information between visible light and thermal infrared images and can leverage the advantages of self-supervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a motion-aware self-supervised RGBT tracking method based on a multi-modal hierarchical Transformer, and belongs to the RGB-T object tracking technology. Background Art

[0002] Object tracking is an important research branch in the field of computer vision, and its goal is to predict the position of the object annotated in the first frame of the video. Although object tracking has made good progress recently, it still faces many challenges such as occlusion, low light, and large-scale changes.

[0003] Fusing visible light and thermal infrared images at the feature level in RGBT tracking is a common operation. The simplest fusion methods include direct summation or treating the thermal infrared channel as the fourth channel and concatenating it with the visible light channels. However, this unweighted summation and concatenation method does not consider the information differences between visible light and thermal infrared images at the feature level. Subsequently, various fusion methods based on convolutional neural networks have been developed. However, none of these methods consider the long-range dependencies between feature channels. With the emergence of the Transformer, Transformer-based fusion tracking methods have become more and more common. However, in the field of RGBT tracking, the feature methods based on Transformer fusion are computationally intensive and mainly rely on spatial features. Therefore, we propose a module for multi-modal hierarchical Transformer feature fusion, called MHTF. This module aims to enhance and fuse the features extracted by the CNN, capture the long-range dependencies between channels of the two modalities, and achieve more precise fusion.

[0004] Currently, some trackers use motion cues to solve problems such as occlusion and fast motion; however, most methods use optical flow or long short-term memory (LSTM) techniques and only utilize the motion information between individual frames in the image. In addition, the supervised methods adopted by most RGBT trackers require a large amount of time-consuming and laborious manual annotation; with the increase in the scale of the RGBT dataset, self-supervised RGBT tracking has become more and more popular. Therefore, improving the solutions to problems such as occlusion and fast motion and integrating self-supervised RGBT tracking methods will improve the effect of object tracking. Summary of the Invention

[0005] Objective of the Invention: Currently, fully supervised trackers face a series of challenges such as occlusion, fast object movement, and large-scale transformation. Since self-supervised object trackers do not have accurate manual annotations and cannot obtain features comparable to those of fully supervised trackers, the obtained features lack robustness. To overcome these deficiencies, the present invention provides a motion-aware self-supervised RGBT tracking method based on a multi-modal hierarchical Transformer. By adding an MAM module to help the self-supervised RGBT tracker overcome the challenge of fast object movement, and by designing an MHTF module to fully utilize the complementary information between visible light and infrared features, so as to effectively capture the long-distance dependencies between the two modal features on the channel and achieve accurate fusion, and finally obtain enhanced fused features. By introducing a classification enhancement score map based on a multi-head cross-attention mechanism on the basis of the CNN classification head to assist in achieving more accurate classification.

[0006] Technical Solution: To achieve the above objective, the technical solution adopted by the present invention is as follows:

[0007] A motion-aware self-supervised RGBT tracking method based on a multi-modal hierarchical Transformer, comprising the following steps:

[0008] (1) Respectively input the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z i into a feature extraction network for feature extraction to obtain the corresponding visible light search frame feature thermal infrared search frame feature visible light template frame feature and thermal infrared template frame feature The feature extraction network is composed of the first three layers of the ResNet50 network;

[0009] (2) Use a multi-modal hierarchical Transformer feature fusion module (MHTF module) to fuse cross-modal features. The visible light search frame feature and the thermal infrared search frame feature are fused to obtain a search frame fusion feature The visible light template frame feature and the thermal infrared template frame feature are fused to obtain a template frame fusion feature

[0010] (3) Use in the template frame fusion feature as the template feature, use in the search frame fusion feature as the search feature, and use the search feature The search features of it and the subsequent N sampling points are input into the Motion Awareness Module (MAM module) to calculate the motion loss

[0011] (4) The fused features and are input into the convolutional module for convolution-based cross-correlation operation and calculate the classification loss and the regression loss The fused features and are input into the multi-head cross-attention module for multi-head cross-attention-based cross-correlation operation and calculate the classification enhancement loss According to the classification loss the regression loss and the classification enhancement loss calculate the offline loss

[0012] (5) According to the motion loss and the offline loss calculate the total loss Use the total loss to train the network model in stages, and finally obtain the trained network model;

[0013] (6) Use the trained network model to track the video frames to obtain the tracking results.

[0014] Specifically, in the step (1), the RGB image of the video frame and the corresponding thermal infrared image are respectively used as the visible light search frame x v and the thermal infrared search frame x i , use the saliency object detection and dynamic programming method in the unsupervised single-object tracker USOT to obtain the pseudo-labels of the video frames, and use the pseudo-labels to crop the visible light search frame x v and the thermal infrared search frame x i to obtain the visible light template frame z v and the thermal infrared template frame z i , take x = (x v , x i ) as the search frame, and take z = (z v , z i ) as the template frame; respectively obtain the image features of the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z i through the feature extraction network, and obtain the corresponding visible light search frame features thermal infrared search frame features visible light template frame features and thermal infrared template frame features

[0015] Specifically, in step (2), a multi-modal hierarchical Transformer feature fusion module is used to fuse cross-modal features;

[0016] (21) Fuse visible light search frame features with thermal infrared search frame features to obtain search frame fusion features including the following steps:

[0017] (211) First, concatenate and together in the channel dimension, and then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, and W x represents the width of the image feature and H x represents the height of the image feature ;

[0018] (212) Perform a 1×1 convolution operation on and respectively to reduce the channel dimension from 1024 to 256, maintaining the same channel dimension as the concatenation result, obtaining

[0019] (213) Flatten and in the spatial dimension to obtain and

[0020] (214) Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain MultiHead represents the multi-head cross-attention operation;

[0021] (215) Perform a non-linear layer operation on and respectively, expressed as:

[0022]

[0023]

[0024] Among them: Norm represents layer normalization, and FFN represents a feed-forward neural network;

[0025] (216) Perform a non-linear layer operation on and which is expressed as:

[0026]

[0027] (217) Calculate the fused feature of the search frame

[0028] (22) Fuse the visible light template frame feature and the thermal infrared template frame feature to obtain the fused feature of the template frame including the following steps:

[0029] (221) First, concatenate and together in the channel dimension, and then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, and W z represents the width of the image feature and H z represents the height of the image feature ;

[0030] (222) Perform a 1×1 convolution operation on and respectively to reduce the channel dimension from 1024 to 256, maintaining the same channel dimension as the concatenation result, obtaining

[0031] (223) Flatten and in the spatial dimension to obtain and

[0032] (224) Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain MultiHead represents the multi-head cross-attention operation;

[0033] (225) Perform a non - linear layer operation on and respectively, which is expressed as:

[0034]

[0035]

[0036] where: Norm represents layer normalization, and FFN represents a feed - forward neural network;

[0037] (226) Perform a non - linear layer operation on and respectively, which is expressed as:

[0038]

[0039] (227) Calculate the fused feature of the template frame

[0040] Specifically, in step (3), take the F in the fused feature of the template frame z 0 as the template feature, take the in the fused feature of the search frame as the search feature, and form a memory frame feature sequence with the search features of the N sampling points after the search feature Input it into the motion perception module, calculate the average motion vector of the inter - frame coordinates, and infer the linear relationship of the motion vectors through the inter - frame interval, including the following steps: Input it into the motion perception module, calculate the average motion vector of the inter - frame coordinates, and infer the linear relationship of the motion vectors through the inter - frame interval, including the following steps:

[0041] (31) Use the multi - head cross - attention mechanism to calculate the attention weight Attn z 0 between the template feature F and the search feature 0 :

[0042]

[0043] where: Attn 0 represents the attention weight between the template feature and the search feature , represents transpose, represents the scaling factor, and softmax represents the normalization operation;

[0044] (32) Use the multi - head cross - attention mechanism to calculate the attention weight Attn and between the template feature n :

[0045]

[0046] Among them: Attn n represents the template feature and the attention weight between them, represents the transpose of, n = 1, 2,..., N, represents the nth frame in the memory frame feature sequence, which is also the search feature the search feature at the nth sampling point after;

[0047] (33) Use Attn n to weight the coordinate mapping C of the template feature z , so as to obtain and the search feature the inter-frame interval scroll M n→0 :

[0048] M n→0 = Attn n × C z - C x

[0049] Among them: represents the coordinate mapping of the template feature F z 0 ; represents the coordinate mapping of the search feature ;

[0050] (34) According to the assumption of local linear motion, through the inter-frame interval scroll M n→0 obtain a motion estimate of an inter-frame interval:

[0051] M n→(n-1) = M n→0 - M (n-1)→0

[0052] Among them: M n→(n-1) represents the motion estimate of the inter-frame interval between the nth frame and the (n - 1)th frame in the memory frame feature sequence, and there are a total of N motion estimates of inter-frame intervals;

[0053] (35) Take the mean of the N motion estimates of inter-frame intervals and perform a linear layer projection to obtain a single inter-frame interval motion vector F Motion :

[0054]

[0055] (36) Predict the search feature through the attention weight and the single inter-frame interval motion vector Coordinate mapping:

[0056] C x = Attn 0 × C z ,C x pre = C x + F Motion

[0057] Where: C x represents the coordinate mapping of the search feature ,C x represents the coordinate mapping of the search feature generated by using the multi-head cross-attention mechanism ,C x pre represents the predicted value of the coordinate mapping of the search feature obtained by using motion estimation;

[0058] (37) Use the motion loss to constrain the consistency between C x and C x pre .

[0059] Specifically, in the step (4), the fused feature and are input into the convolution module for convolution-based cross-correlation operation, and the classification loss and the regression loss The fused feature and are input into the multi-head cross-attention module for multi-head cross-attention-based cross-correlation operation, and the classification enhancement loss includes the following steps:

[0060] (41) Take and as inputs, and use the cross-entropy loss as the classification loss Use the IoU loss as the regression loss The classification enhancement score map adopted is expressed as

[0061] (42) The cross-entropy loss uses the classification enhancement score map to calculate the classification enhancement loss for calculating the classification loss The label size is 25×25, and for calculating the classification enhancement loss the label size is 31×31;

[0062] (43) According to the classification loss the classification enhancement loss ​And regression loss Calculate the offline loss Where: λ is a hyperparameter;

[0063] (44) Calculate the total loss Where: β is a hyperparameter.

[0064] Specifically, in step (5), using the total loss Perform staged tracking training on the network model, and finally obtain the trained network module. Requirements for staged training: After importing the ImageNet dataset into the pre-trained feature extraction network, fix the parameters of the feature extraction network, and add the motion loss after the sixth training stage Make the feature extraction network trainable after the tenth training stage, and finally obtain all the parameters of the feature extraction network.

[0065] Specifically, in step (6), use the trained network model to track the video frames to obtain the tracking results. Requirements for tracking:

[0066] ① The motion perception module does not participate in the tracking process of the trained network model;

[0067] ② Online template update allows adaptation to appearance changes, but may accumulate errors during tracking failure, resulting in incorrect tracking; Therefore, to ensure the correctness of the template and enhance the efficiency of tracking, the template is not updated during the tracking process, that is, during the tracking process, no online template is used, and the visible light template frame z v and the thermal infrared template frame z i are not updated;

[0068] ③ The tracking result is represented as the final classification score map Where: represents the final classification score map, represents the original classification score map obtained by inputting the template feature and the search feature into the convolutional module, represents the classification enhancement score map obtained by inputting the template feature F z 0 and the search feature into the multi-head cross-attention module, is used to adjust the classification enhancement score map to be the same size as the original classification score map ;

[0069] ④ Use the maximum value of the final classification score map as the center position of the tracking target, and perform network regression to determine the position of the bounding box.

[0070] Beneficial effects: The motion-aware self-supervised RGB-T tracking method based on multi-modal hierarchical Transformer provided by the present invention has the following advantages compared with the prior art: (1) The present invention can make full use of the complementary information between visible light and thermal infrared images without introducing excessive computational complexity, and give play to the advantage of Transformer in capturing long-range dependencies; (2) The MHTF module in the present invention adopts two-layer non-linear layer operations, which can perform more in-depth non-linear processing on features; (3) Compared with the defect that a large number of labeled images are required for a fully supervised model and a large amount of manpower and material resources are needed, the present invention adopts a self-supervised training method to overcome this defect, and there are fewer RGB-T trackers using the self-supervised training method, which is beneficial to the further development of self-supervised RGB-T object tracking technology. Brief Description of the Drawings

[0071] Figure 1 is the implementation flowchart of the method of the present invention;

[0072] Figure 2 is the structural schematic diagram of the S2OTFormer framework for implementing the method of the present invention;

[0073] Figure 3 is the structural block diagram of the MAM module. Detailed Embodiment

[0074] The present invention will be specifically introduced below in conjunction with the accompanying drawings and specific embodiments.

[0075] As Figure 1 shown is the implementation flowchart of a motion-aware self-supervised RGB-T tracking method based on multi-modal hierarchical Transformer. First, a self-supervised method is used to generate pseudo-labels, and then ResNet50 is used to extract the features of RGB images and thermal infrared images; then the MHTF module is used to capture the long-range dependencies between the two modal features on the channel to achieve more accurate fusion; then the cross-correlation operation based on convolution is performed on the fused features, and then the classification-enhanced score map based on the multi-head cross-attention mechanism is used to assist in achieving accurate classification; the MAM module is introduced to record the search frame features and extract the corresponding motion vectors, and these vectors are used to strengthen the consistency with the current search frame features during the training of the network model; then, the losses obtained from the cross-correlation operation and the MAM module are minimized; finally, the test video frames are input into the trained network for tracking to obtain the response map, which is the predicted target position. Each step will be specifically described below.

[0076] Step S01: Respectively, the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z iThe input feature extraction network extracts features to obtain the corresponding visible light search frame features Thermal infrared search frame features Visible light template frame features And thermal infrared template frame features

[0077] The feature extraction network consists of the first three layers of the ResNet50 network.

[0078] Take the RGB image of the video frame and the corresponding thermal infrared image as the visible light search frame x v And the thermal infrared search frame x i , use the saliency object detection and dynamic programming method in the unsupervised single-object tracker USOT to obtain the pseudo-labels of the video frames, and use the pseudo-labels to respectively process the visible light search frame x v And the thermal infrared search frame x i Perform cropping to obtain the visible light template frame z v And the thermal infrared template frame z i , take x = (x v , x i ) as the search frame, and take z = (z v , z i ) as the template frame.

[0079] Step S02: Use the MHTF module to fuse cross-modal features to obtain the search frame fusion features And the template frame fusion features

[0080] Abbreviate the multi-modal hierarchical Transformer feature fusion module as the MHTF module. The process of the MHTF module is as Figure 2 Shown.

[0081] Step 21: Fuse the visible light search frame features With the thermal infrared search frame features To obtain the search frame fusion features The steps are as follows:

[0082] Step 211: First, concatenate And Together in the channel dimension, and then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, and W x Represents the image feature Of the width, and H x Represents the image feature Of the height;

[0083] Step 212: Perform a 1×1 convolution operation on and respectively, reducing the channel dimension from 1024 to 256 and maintaining the same channel dimension as the concatenation result to obtain

[0084] Step 213: Flatten and in the spatial dimension to obtain and

[0085] Step 214: Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain Use as the query vector q, as the key vector k and value vector v, and perform a multi-head cross-attention operation to obtain MultiHead represents the multi-head cross-attention operation;

[0086] Step 215: Perform a non-linear layer operation on and respectively, expressed as:

[0087]

[0088]

[0089] where: Norm represents layer normalization, and FFN represents the feed-forward neural network;

[0090] Step 216: Perform a non-linear layer operation on and respectively, expressed as:

[0091]

[0092] Step 217: Calculate the search frame fusion feature through to reduce information loss. Step 22: Fuse the visible light template frame feature f z v and the thermal infrared template frame feature to obtain the template frame fusion feature including the following steps:

[0093] Step 221: First, on the channel dimension, and Connected together, then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, W z represents the image feature f z vi width, H z represents the image feature f z vi height;

[0094] Step 222: Respectively perform a 1×1 convolution operation on and to reduce the channel dimension from 1024 to 256, maintaining the same channel dimension as the concatenation result, obtaining

[0095] Step 223: Flatten and in the spatial dimension, obtaining and

[0096] Step 224: Use as the query vector q, as the key vector k and value vector v, perform the multi-head cross-attention operation, obtaining Use as the query vector q, as the key vector k and value vector v, perform the multi-head cross-attention operation, obtaining MultiHead represents the multi-head cross-attention operation;

[0097] Step 225: Respectively perform a non-linear layer operation on and which is expressed as:

[0098]

[0099]

[0100] where: Norm represents layer normalization, FFN represents the feed-forward neural network;

[0101] Step 226: Perform a non-linear layer operation on and which is expressed as:

[0102]

[0103] Step 227: Calculate the template frame fusion feature through Reduce information loss.

[0104] Step S03: Calculate the motion loss through the MAM module

[0105] Abbreviate the motion perception module as the MAM module. As Figure 3 shown, take in the template frame fusion feature as the template feature, and take in the search frame fusion feature as the search feature, and form a memory frame feature sequence with the search features of the next N sampling points of the search feature Input it into the motion perception module, calculate the average motion vector of the inter-frame coordinates, and infer the linear relationship of the motion vector through the inter-frame interval, including the following steps:

[0106] Step 31: Use the multi-head cross-attention mechanism to calculate the attention weight Attn z 0 between the template feature F and the search feature 0 :

[0107]

[0108] where: Attn 0 represents the attention weight between the template feature and the search feature , represents transpose of represents the scaling factor, and softmax represents the normalization operation;

[0109] Step 32: Use the multi-head cross-attention mechanism to calculate the attention weight Attn and between n :

[0110]

[0111] where: Attn n represents the attention weight between the template feature and , represents transpose of, n = 1, 2,..., N, represents the nth frame in the memory frame feature sequence, which is also the search feature of the nth sampling point after the search feature

[0112] Step 33: Use Attn n to weight the template feature Coordinate mapping C of z , thus obtaining and search feature The inter-frame interval scroll M between n→0 :

[0113] M n→0 = Attn n × C z - C x

[0114] Where: Represents the coordinate mapping of the template feature of, Represents the coordinate mapping of the search feature of;

[0115] Step 34: According to the assumption of local linear motion, through the inter-frame interval scroll M n→0 Obtain a motion estimate of one inter-frame interval:

[0116] M n→(n-1) = M n→0 - M (n-1)→0

[0117] Where: M n→(n-1) Represents the motion estimate of the inter-frame interval between the nth frame and the (n - 1)th frame in the memory frame feature sequence, and there are a total of N motion estimates of inter-frame intervals;

[0118] Step 35: Calculate the mean of the N motion estimates of inter-frame intervals and perform a linear layer projection to obtain a single inter-frame interval motion vector F Motion :

[0119]

[0120] Step 36: Predict the coordinate mapping of the search feature through the attention weight and the single inter-frame interval motion vector:

[0121] C x = Attn 0 × C z , C x pre = C x + F Motion

[0122] Where: C x Represents the coordinate mapping of the search feature of, C x Represents the coordinate mapping of the search feature F x 0 generated by the multi-head cross-attention mechanism, C xpre Indicates the search features obtained by motion estimation The predicted value of coordinate mapping;

[0123] Step 37: Use the motion loss to constrain C x and C x pre for consistency between them.

[0124] Step S04: Calculate the offline loss according to the classification loss regression loss classification enhancement loss Calculate the offline loss

[0125] The fused features are correlated with the input convolution module through convolution-based cross-correlation operation, and calculate the classification loss and regression loss The fused features are correlated with the input multi-head cross-attention module through multi-head cross-attention-based cross-correlation operation, and calculate the classification enhancement loss According to the classification loss regression loss and classification enhancement loss calculate the offline loss

[0126] Step 41: Take and as inputs, use the cross-entropy loss as the classification loss Use the IoU loss as the regression loss The classification enhancement score map adopted is expressed as

[0127] Step 42: The cross-entropy loss uses the classification enhancement score map to calculate the classification enhancement loss used to calculate the classification loss The label size for calculating the classification loss is 25×25, and the label size for calculating the classification enhancement loss is 31×31; due to the different label sizes used, the classification enhancement score map can introduce more scale information into the network model;

[0128] Step 43: According to the classification loss classification enhancement loss and regression loss calculate the offline loss where: λ is a hyperparameter, and in this case, λ = 0.2;

[0129] Step 44: Calculate the total loss Where: β is a hyperparameter, and in this case, β = 0.2 is taken.

[0130] Step S05: According to the motion loss and the offline loss Calculate the total loss Use the total loss Perform staged training on the network model, and finally obtain a trained network model.

[0131] There are the following requirements during staged training: After importing the ImageNet dataset into the pre-trained feature extraction network, fix the parameters of the feature extraction network, and add the motion loss after the sixth training stage Make the feature extraction network trainable after the tenth training stage, and finally obtain all the parameters of the feature extraction network.

[0132] Step S06: Use the trained network model to track the video frames to obtain the tracking results.

[0133] There are the following requirements during tracking:

[0134] ① The motion perception module does not participate in the tracking process of the trained network model;

[0135] ② Online template update allows adaptation to appearance changes, but errors may accumulate during tracking failures, resulting in incorrect tracking; therefore, to ensure the correctness of the template and enhance the efficiency of tracking, the template is not updated during the tracking process, that is, during the tracking process, no online template is used, and the visible light template frame z v and the thermal infrared template frame z i are not updated.

[0136] ③ The tracking result is represented as the final classification score map Where: represents the final classification score map, represents the original classification score map obtained by inputting the template feature and the search feature into the convolutional module ( Figure 2 the Conv module in represents the classification enhancement score map obtained by inputting the template feature F z 0 and the search feature into the multi-head cross-attention module ( Figure 2 the Multi-Head Attention module in is used to adjust the classification enhancement score map to be the same as the original classification score map of the same size;

[0137] ④ Use the maximum value of the final classification score map as the center position of the tracking target, and perform network regression to determine the position of the bounding box.

[0138] A method for implementing the above-mentioned multi-modal hierarchical Transformer-based motion perception self-supervised RGBT tracking method, adopting the S2OTFormer framework, mainly including a feature extraction network, an MHTF module, an MAM module, a correlation network, and a tracking head. The tracking head includes a common classification head, a common regression head, and a classification enhancement head based on the multi-head cross attention mechanism; the entire S2OTFormer framework adopts a Siamese network architecture.

[0139] The feature extraction network is composed of the first three layers of the ResNet50 network, and the input is the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z i , and the output is the visible light search frame feature the thermal infrared search frame feature the visible light template frame feature and the thermal infrared template frame feature

[0140] The MHTF module performs feature enhancement and fusion on the visible light search frame feature and the thermal infrared search frame feature to generate a search frame fusion feature The MHTF module performs feature enhancement and fusion on the visible light template frame feature and the thermal infrared template frame feature to generate a template frame fusion feature

[0141] The search frame fusion feature and the template frame fusion feature perform feature interaction through the correlation network. The correlation network includes a convolutional module and a multi-head cross attention module; the fused feature and are input into the convolutional module for convolution-based cross-correlation operation, and the classification loss and the regression loss The fused feature and F z 0 are input into the multi-head cross attention module for multi-head cross attention-based cross-correlation operation, and the classification enhancement loss According to the classification loss Regression loss and classification enhancement loss Calculate the offline loss

[0142] Among the fused features of the template frame As the template feature, among the fused features of the search frame As the search feature, and the search feature The search features of the subsequent N sampling points form the memory frame feature sequence Input the MAM module, calculate the average motion vector of the inter-frame coordinates, infer the linear relationship of the motion vector through the inter-frame interval, and then calculate the motion loss Furthermore, calculate the total loss Use the total loss to train the network model in stages, and finally obtain the trained network model.

[0143] The process of using the trained network model to track video frames is basically the same as the training process. First, the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z i Input the feature extraction network to obtain the visible light search frame feature Thermal infrared search frame feature Visible light template frame feature and thermal infrared template frame feature Then, the visible light search frame feature Thermal infrared search frame feature Visible light template frame feature and thermal infrared template frame feature Input the MHTF module to obtain the fused feature of the search frame and the fused feature of the template frame Next, and The original classification score map obtained by performing cross-correlation operation based on convolution by inputting into the convolution module Use the template feature and the search feature Input the multi-head cross-attention module to perform cross-correlation operation based on multi-head cross-attention to obtain the classification enhancement score map Finally, calculate the final classification score map Among them, the maximum value of the final classification score map Is used as the center position of the tracking target, and network regression is performed to determine the position of the bounding box.

[0144] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and any technical solutions obtained by means of equivalent substitution or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A motion-aware self-supervised RGBT tracking method based on a multi-modal hierarchical Transformer, characterized in that: It includes the following steps: (1) Visible light search frame x is separately v , thermal infrared search frame x i , visible light template frame z v and thermal infrared template frame z i are input into the feature extraction network for feature extraction to obtain the corresponding visible light search frame features , thermal infrared search frame features , visible light template frame features and thermal infrared template frame features The feature extraction network is composed of the first three layers of the ResNet50 network; (2) Use the multi-modal hierarchical Transformer feature fusion module to fuse cross-modal features, including visible light search frame features and thermal infrared search frame features After fusion, the search frame fusion features are obtained Visible light template frame features and thermal infrared template frame features After fusion, the template frame fusion features are obtained (3) Take the in the template frame fusion feature as the template feature, and take the in the search frame fusion feature as the search feature. Input the search feature and the search features of the next N sampling points into the motion perception module to calculate the motion loss (4) The fused features and are input into a convolutional module for convolution-based cross-correlation operations, and the classification loss and regression loss The fused features and are input into a multi-head cross-attention module for multi-head cross-attention-based cross-correlation operations, and the classification enhancement loss Based on the classification loss regression loss and classification enhancement loss calculate the offline loss including the following steps: (41) Taking and as inputs, using cross-entropy loss as the classification loss using IoU loss as the regression loss The classification enhancement score map adopted is represented as (42) Cross-entropy loss utilizes classification-enhanced score maps Calculate classification-enhanced loss Used to calculate classification loss The label size for calculating classification-enhanced loss is 25×25 The label size is 31×31; (43)According to the classification loss Classification enhancement loss and regression loss Calculate the offline loss where: λ is a hyperparameter; (44) Calculate the total loss where: β is a hyperparameter; Calculate the total loss according to the motion loss and the offline loss Calculate the total loss Use the total loss Perform staged training on the network model to finally obtain a trained network model; (6) Use the trained network model to track the video frames to obtain the tracking results.

2. The motion-aware self-supervised RGBT tracking method based on the multi-modal hierarchical Transformer according to claim 1, characterized in that: In the said step (1), the RGB image of the video frame and the corresponding thermal infrared image are respectively used as the visible light search frame x v and the thermal infrared search frame x i . The pseudo - labels of the video frame are obtained by using the saliency object detection and dynamic programming method in the unsupervised single - object tracker USOT. The visible light search frame x v and the thermal infrared search frame x i are respectively cropped to obtain the visible light template frame z v and the thermal infrared template frame z i . Let x=(x v ,x i ) be the search frame, and z=(z v ,z i ) be the template frame. The image features of the visible light search frame x v , the thermal infrared search frame x i , the visible light template frame z v and the thermal infrared template frame z i are respectively obtained through the feature extraction network, and the corresponding visible light search frame feature thermal infrared search frame feature visible light template frame feature and thermal infrared template frame feature 3. The motion-aware self-supervised RGBT tracking method based on the multi-modal hierarchical Transformer according to claim 2, characterized in that: In the step (2), a multi-modal hierarchical Transformer feature fusion module is used to fuse cross-modal features; (21)Fuse the visible light search frame features with the thermal infrared search frame features to obtain the search frame fusion features The steps are as follows: (211) First, concatenate and together in the channel dimension, and then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, and W x represents the width of the image feature , and H x represents the height of the image feature ; (212) Perform a 1×1 convolution operation on and respectively, reducing the channel dimension from 1024 to 256 to obtain (213) Flatten and in the spatial dimension to obtain and (214) Take as the query vector q, as the key vector k and value vector v, perform the multi-head cross-attention operation to obtain Take as the query vector q, as the key vector k and value vector v, perform the multi-head cross-attention operation to obtain MultiHead represents the multi-head cross-attention operation; (215) Perform a non-linear layer operation on and respectively, which is expressed as: Where: Norm represents layer normalization, and FFN represents a feed-forward neural network; (216) pair and perform a non-linear layer operation, expressed as: (217) Calculate the fused features of the search frame (22)Fuse the visible light template frame features and the thermal infrared template frame features to obtain the template frame fusion features The steps are as follows: (221) First, concatenate and together in the channel dimension, and then perform a 1×1 convolution operation to reduce the channel dimension from 2048 to 256, obtaining Concat represents the feature concatenation operation, Proj represents the 1×1 convolution operation, W z represents the width of the image feature and H z represents the height of the image feature ; (222) Perform a 1×1 convolution operation on and respectively, reducing the channel dimension from 1024 to 256 to obtain (223) Flatten and in the spatial dimension to obtain and (224) Take as the query vector q, as the key vector k and value vector v, perform multi-head cross-attention operation to obtain Take as the query vector q, as the key vector k and value vector v, perform multi-head cross-attention operation to obtain MultiHead represents the multi-head cross-attention operation; (225)Perform a non - linear layer operation on and respectively, which is expressed as: Where: Norm represents layer normalization, and FFN represents a feed-forward neural network; (226)Pair and perform a non-linear layer operation, expressed as: (227) Calculate the fused features of the template frames 4. The motion perception self-supervised RGBT tracking method based on the multi-modal hierarchical Transformer according to claim 3, characterized in that: In the step (3), take the in the fused feature of the template frame as the template feature, take the in the fused feature of the search frame as the search feature, and form a memory frame feature sequence with the search features of the subsequent N sampling points of the search feature Input it into the motion perception module, calculate the average motion vector of the inter-frame coordinates, and infer the linear relationship of the motion vector through the inter-frame interval, including the following steps: (31) Calculate the template features using the multi-head cross-attention mechanism and the search features The attention weight Attn between 0 : Where: Attn 0 represents the template feature and the attention weight between the search features , represents the transpose of represents the scaling factor, and softmax represents the normalization operation; (32) Calculate the template features using the multi-head cross-attention mechanism and the attention weight Attn between n : Where: Attn n represents the template feature and the attention weight between represents the transpose of, where n = 1, 2, …, N represents the n-th frame in the memory frame feature sequence, which is also the search feature and the search feature at the n-th sampling point after (33) Use Attn n to weight the template features of the coordinate mapping C z , thereby obtaining the inter-frame interval scroll M between and the search features n→0 : M n→0 = Attn n × C z - C x Wherein: represents the coordinate mapping of the template feature ; represents the coordinate mapping of the search feature ; (34) According to the assumption of local linear motion, roll M through the inter-frame interval n→0 to obtain the motion estimation of an inter-frame interval: M n→(n-1) = M n→0 -M (n-1)→0 Where: M n→(n-1) represents the motion estimation of the inter-frame interval between the nth frame and the (n - 1)th frame in the memory frame feature sequence, and there are a total of N motion estimations of inter-frame intervals; (35) Average the motion estimations for N inter-frame intervals and perform a linear layer projection to obtain a single inter-frame interval motion vector F Motion : (36)Predicting search features through attention weights and single inter-frame interval motion vectors Coordinate mapping of C x = Attn 0 × C z , C x pre = C x + F Motion Among them: C x represents the coordinate mapping of the search feature , C x represents the coordinate mapping of the search feature generated by using the multi-head cross-attention mechanism , C x pre represents the predicted value of the coordinate mapping of the search feature obtained by using motion estimation ; (37) Use motion loss to constrain C x and C x pre for consistency between them.

5. The motion-aware self-supervised RGBT tracking method based on the multi-modal hierarchical Transformer according to claim 4, characterized in that: In the step (5), using the total loss to perform phased tracking training on the network model, and finally obtaining a trained network module. During phased training, the requirements are as follows: after importing the ImageNet dataset into the pre-trained feature extraction network, fix the parameters of the feature extraction network, and add the motion loss after the sixth training phase make the feature extraction network trainable after the tenth training phase, and finally obtain all the parameters of the feature extraction network.

6. The motion-aware self-supervised RGBT tracking method based on the multi-modal hierarchical Transformer according to claim 5, wherein: In the step (6), use the trained network model to track the video frames to obtain the tracking results. When tracking, the requirements are: ① The motion perception module does not participate in the tracking process of the trained network model; ② During the tracking process, no online template is used, and the visible light template frame z v and the thermal infrared template frame z i are not updated; ③ The tracking result is represented as the final classification score map Wherein: represents the final classification score map, represents the original classification score map obtained by inputting the template feature and the search feature into the convolutional module; represents the classification enhancement score map obtained by inputting the template feature and the search feature into the multi-head cross-attention module, which is used to adjust the classification enhancement score map to be the same size as the original classification score map ; ④Use the maximum value of the final classification score map as the center position of the tracking target, and perform network regression to determine the position of the bounding box.

Citation Information

Cited By

  • An unmanned aerial vehicle tracking method and system based on multi-feature fusion

    CN122760636A