An RGBT Object Tracking Method Based on Convolutional Attention Fusion
By inserting a convolutional attention fusion module between the Transformer encoder, the local feature utilization and cross-modal feature fusion of the RGBT target tracking method are improved, and the problem of insufficient robustness and accuracy in the existing methods is solved, achieving a more stable target tracking effect.
Patent Information
- Application Number
- CN202510582564.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing RGBT target tracking method fails to effectively utilize local and cross-modal features when fusing visible light and thermal infrared modal information, resulting in insufficient robustness and accuracy.
A RGBT target tracking method based on convolutional attention fusion is designed. By inserting a convolutional attention fusion module between the Transformer encoder, the utilization of local features and cross-modal feature fusion are improved, and attention to global information is maintained.
A more stable RGBT target tracking performance is achieved, the model's utilization of visible light and thermal infrared modal features is improved, and the tracking is enhanced.
Smart Images

Figure CN120088292B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to an RGBT object tracking method based on convolutional attention fusion. Background Art
[0002] Object tracking is one of the important research contents in the field of computer vision, aiming to continuously track an object in a given video sequence, and it has been widely used in many fields such as video surveillance, human-computer interaction, and visual navigation. Early tracking algorithms tended to use handcrafted features to track objects based on domain knowledge and experience. However, the performance of these methods is difficult to meet the requirements in real scenarios. Since the proposal of deep learning, with its powerful feature modeling ability, it has gradually replaced traditional algorithms and become the mainstream method in the field of object tracking.
[0003] In recent years, with the proposal of deep learning models such as convolutional neural networks and Transformers, object tracking methods based on deep learning have made significant progress. These methods learn image features in a data-driven manner and improve the performance of object tracking. RGB object tracking only uses the visible light modality to track objects and still faces many challenges. Visible light-based tracking methods are vulnerable to factors such as light intensity, object color, and weather conditions. Some researchers have tried to introduce data from other modalities and fuse it with visible light data to improve the accuracy of the tracking algorithm using the complementary information of different modalities. Infrared light has a thermal effect, can reflect the temperature of an object, and is not affected by light, and can identify objects at night. Therefore, infrared light has broad application prospects in the field of object tracking. RGBT object tracking comprehensively uses thermal infrared information and visible light information to perform object tracking. However, thermal infrared images have defects such as severe noise, low resolution, unclear objects, insignificant texture features, and difficulty in penetrating transparent objects. In contrast, although visible light provides rich texture information and color information, it is easily affected by lighting conditions.
[0004] In the field of RGBT object tracking, how to effectively fuse the information of the visible light and thermal infrared modalities and utilize their respective advantages to improve the robustness and accuracy of tracking is a particularly important research point. Existing RGBT object tracking methods can be divided into pure convolutional-based methods, convolutional-Transformer hybrid methods, and pure Transformer-based methods. However, most methods still focus on the global fusion of modalities and fail to consider the fusion and utilization of local features, nor fully exploit the potential of convolution and the attention mechanism.
[0005] Based on this, the present invention designs an RGBT object tracking method based on convolutional attention fusion to solve the above problems. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the present invention provides an RGBT object tracking method based on convolutional attention fusion. By designing a convolutional attention fusion module, the tracker's utilization of local features is improved, cross-modal feature fusion is promoted, and attention to global information is maintained, thereby achieving more stable RGBT object tracking.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0008] An RGBT object tracking method based on convolutional attention fusion, comprising the following steps:
[0009] Step 1, video preprocessing;
[0010] Randomly select a video sequence from the dataset, where each frame is an image; select the rectangular area where the target is located at the same position in the first frame of the visible light and thermal infrared modalities, scale this area, and save it as the target template of the video sequence; starting from the second frame, use the target position of the previous frame as the center point, select a square area with a range larger than the target area, and scale it as the search area of the current frame;
[0011] Step 2, feature extraction;
[0012] Cut the target template and search area of the visible light and thermal infrared modalities into several blocks and expand and splice them. Map the image information into a one-dimensional feature sequence through a linear mapping layer, and add global position encoding to it; use the Transformer encoder of the backbone network with shared parameters to extract features for the features of the two modalities respectively;
[0013] Step 3, local feature enhancement and fusion;
[0014] Insert a convolutional attention fusion module between the Transformer encoders: perform two-dimensional processing on the one-dimensional feature sequence through a sliding window, then perform zero-padding around it, use the sliding window to slide cyclically starting from the upper left corner, the sliding window traverses the entire feature map, and select the local area; serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities; finally, merge the local cross-attention results of each sliding window;
[0015] Step 4, object tracking prediction.
[0016] Furthermore, in Step 1, the LasHeR dataset is used as the dataset.
[0017] Furthermore, the backbone network in Step 2 uses OSTrack as the baseline model and expands it into a two-branch, and makes the original input visible light and thermal infrared features be respectively Xr , X t , the i -th layer Transformer encoder is denoted as Encoder i . The feature extraction process can be formulated as: ;
[0018] where, represents the visible light feature of the i-th layer, represents the thermal infrared feature of the i-th layer, represents the visible light feature of the (i + 1)-th layer, represents the thermal infrared feature of the (i + 1)-th layer.
[0019] Furthermore, convolutional attention fusion modules are inserted into the 4th, 7th, and 10th layers between the Transformer encoders of the backbone network.
[0020] Furthermore, in step three, a local area is selected through a sliding window. The specific steps are as follows:
[0021] For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , it is expanded into a two-dimensional feature and extended to ( H +2× padding )×( W +2× padding ) through zero-padding. Among them, H , W, C are the height, width, and number of channels of the original feature map respectively, padding is the padding size; the expanded area is divided into p×p blocks, and each block is regarded as a feature point with a dimension of P ∈ R p×p×C . Among them, P refers to any feature point; subsequently, a sliding window with a side length of ks starts sliding from the upper left corner of the expanded feature map with a step size of stride , and a local area X left:left+ks,top:top+ks is selected, where ( left , top ) is the upper left corner coordinate, and the sliding window is denoted as X' ∈ R ks×ks×C ; the sliding window slides continuously in a loop, and finally, the local features of all local areas are obtained; among them, N num represents the total number of sliding times, and the calculation formula is:
[0022] 。
[0023] Furthermore, in step three, serialize the local region and calculate the local cross-attention of the visible light and thermal infrared modalities. The specific steps are as follows:
[0024] Define local position encoding E ∈R 1×(ks×ks)×1 , broadcast it to the same dimension as S , and add it to all sliding windows; through the local linear mapping layer, map the local region to obtain the local query Q local and the key-value K local ; through the global linear mapping layer, map the original feature map to the global weight V global ;
[0025] Calculate the local cross-attention of the visible light and thermal infrared modalities using the following formula:
[0026] (a)
[0027] where d k represents the feature channel dimension.
[0028] Furthermore, use the method of calculating the weights of all blocks at one time to obtain V∈R H×W×C , V represents the attention weight value, and slide it on V through the same sliding window to obtain .
[0029] Furthermore, the local cross-attention Attn of the visible light and thermal infrared modalities are calculated independently. When calculating the local cross-attention of the visible light modality, use the V global of the thermal infrared modality to calculate Attn. When calculating the local cross-attention of the thermal infrared modality, use the V global of the visible light modality to calculate Attn.
[0030] Furthermore, in step three, merge the local cross-attention results of each sliding window. The specific steps are as follows:
[0031] Split the local attention result into N num local feature maps to obtain S' ∈ R ks×ks×C , S'It represents the local feature map, restores the local feature map to the position of the original feature map corresponding to the sliding window, calculates the average value of the overlapping parts of the local feature map, adds the local attention results according to the corresponding positions, and calculates the attention average value according to the coverage of the sliding window in each block as the final local attention score.
[0032] Furthermore, step four is specifically as follows: The Transformer encoder in the final layer of the backbone network outputs the features of two modalities, visible light and thermal infrared. They are merged using convolution and input to the prediction head for prediction to obtain the target position.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the Transformer tracking architecture, the present invention integrates the ideas of convolution and attention, designs an efficient convolutional attention fusion module to complete local feature enhancement and cross-modal interaction; through the complementarity with the global features of the ViT network, the RGBT target tracking performance is effectively improved.
[0034] The present invention enhances the robustness of the target tracking task through multi-modal fusion. The model extracts features from both visible light and thermal infrared modalities simultaneously, and through the convolutional attention fusion module, realizes local feature enhancement and cross-modal feature interaction, avoiding the influence of single-modal defects on the tracking performance.
[0035] The present invention can be used for any RGBT video sequence. The proposed convolutional attention fusion module can adjust padding, stride, window size, etc., with strong adjustability, and a better combination can be selected according to specific situations. In addition, the convolutional attention fusion module can be applied to most ViT-based trackers.
[0036] In the present invention, video sequences are randomly selected from the LasHeR dataset, and the target templates are calibrated; the bimodal images are scaled to a unified size and cut into multiple blocks, and the features of each block are encoded using linear mapping and flattened into a one-dimensional feature sequence; the Transformer encoder is iteratively used to extract the sequence features; during the feature extraction process, the convolutional attention fusion module is interspersed. First, the one-dimensional feature sequence is merged into a two-dimensional feature map, local regions are selected on the feature map through a sliding window, and the local features and cross-modal features are enhanced through the attention mechanism; the enhanced multi-modal features are merged through convolution and input to the prediction head for tracking prediction; the performance of the present invention on multiple video tracking datasets is better than other advanced RGBT tracking algorithms.
[0037] The present invention can realize more effective utilization of local features and cross-modal features, and solves the deficiencies of existing RGBT trackers in the utilization of local features and cross-modal interaction. Brief Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0039] Figure 1 It is a comparison chart of conventional attention, windowed attention, and convolutional self-attention. Among them: (a) is conventional attention, which directly calculates attention on the complete feature map. (b) is windowed attention, which divides the feature map into multiple non-overlapping windows, calculates local attention within the windows, and as the number of layers deepens, merges adjacent windows. (c) is convolutional self-attention, which calculates local attention of adjacent regions through overlapping sliding windows.
[0040] Figure 2 It is the overall structure diagram of the RGBT target tracking model of the present invention.
[0041] Figure 3 It is a schematic diagram of convolutional self-attention.
[0042] Figure 4 It is a schematic diagram of convolutional cross-attention.
[0043] Figure 5 It is the comparison result of the challenge attributes between the method of the present invention and other methods.
[0044] Figure 6 It is the comparison result of the accuracy rate between the present invention and other methods on the LasHeR dataset.
[0045] Figure 7 It is the comparison result of the success rate between the present invention and other methods on the LasHeR dataset.
[0046] Figure 8 It is the influence of different window sizes on the tracking performance.
[0047] Figure 9 It is Example 1 of the visualization tracking results of the method of the present invention and other RGBT target tracking methods (the fourth boy in the first column).
[0048] Figure 10 It is Example 2 of the visualization tracking results of the method of the present invention and other RGBT target tracking methods (shooting the ball into the basket three times).
[0049] Figure 11 It is Example 3 of the visualization tracking results of the method of the present invention and other RGBT target tracking methods (the boy playing with the mobile phone). Specific implementation manners
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0051] Embodiment 1: An RGBT target tracking method based on convolutional attention fusion, comprising the following steps:
[0052] Step 1, video preprocessing;
[0053] In a common RGBT target tracking dataset, randomly select a video sequence, where each frame is an image; select the rectangular region where the target is located at the same position in the first frame of the visible light and thermal infrared modalities, scale the region to 128×128 pixels, and save it as the target template of the video sequence; starting from the second frame, with the target position of the previous frame as the center point, assuming the width and height of the region where the target is located in the previous frame are w and h respectively, select a square region with a side length of and scale it to 256×256 pixels as the search region for the current frame;
[0054] Furthermore, in Step 1, select the video sequence in the LasHeR dataset as the tracking video; the LasHeR dataset contains a total of 1224 video sequences and 734.8K frames, with rich data resources and wide coverage of real-world scenarios, which is sufficient for model training.
[0055] Step 2, feature extraction;
[0056] Cut the target template and search region of the visible light and thermal infrared modalities into several 16×16 blocks and expand and splice them. Map the image information into a one-dimensional feature sequence through a linear mapping layer, and add global position encoding to it; use the Transformer encoder of the backbone network with shared parameters to extract features for the features of the two modalities respectively.
[0057] The size of the search region is 256×256, which is cut into blocks of size 16×16 in Step 2, a total of 16×16 blocks. After passing through the linear mapping layer, the block information is mapped to a point to obtain a feature map. Therefore, the size of the feature map is 16×16, and each feature point in it corresponds to the original 16×16 block.
[0058] Furthermore, in Step 2, the backbone network uses OSTrack as the baseline model and is extended to a two-branch structure, making the original input visible light and thermal infrared features Xr , X t , the i -th layer Transformer encoder is denoted as Encoder i . The feature extraction process can be formulated as: ;
[0059] where, represents the visible light feature of the i-th layer, represents the thermal infrared feature of the i-th layer, represents the visible light feature of the (i + 1)-th layer, represents the thermal infrared feature of the (i + 1)-th layer.
[0060] Step 3: Local feature enhancement and fusion;
[0061] Insert a convolutional attention fusion module at specified positions (the 4th, 7th, and 10th layers) between the Transformer encoders of the backbone network. The steps are as follows: Two-dimensionally process the one-dimensional feature sequence through a sliding window, then perform zero-padding around it, use the sliding window to slide cyclically starting from the upper left corner, and the sliding window traverses the entire feature map to select a local area; Serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities; Finally, merge the local cross-attention results of each sliding window to enhance local information and cross-modal feature interaction.
[0062] The convolutional attention fusion module is implemented using convolutional cross-attention. The convolutional attention fusion module, as a plug-and-play module, can be embedded into most trackers using ViT and is applicable to both single-modal and dual-modal cases. The inventive method can achieve relatively stable tracking for any RGBT video sequence and has high performance in most challenging scenarios.
[0063] Furthermore, in Step 3, the local area is selected through a sliding window, and the specific steps are as follows:
[0064] For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , expand it into a two-dimensional feature and expand it to ( H +2× padding )×( W +2× padding ) through zero-padding, where H , W, C are the height, width, and number of channels of the original feature map respectively, padding is the padding size; Divide the expanded area into p×p blocks, and regard each block as a block with a dimension of P∈ R p×p×C feature points, where P represents any feature point; subsequently, a sliding window with a side length of ks (such as a sliding window with a size of 2×2) starts sliding from the upper left corner of the expanded feature map, with a step size of stride (such as 1), and a local area X left:left+ks,top:top+ks is selected, where ( left , top ) are the upper left coordinates, and the sliding window is represented as X' ∈ R ks×ks×C ; by continuously sliding this sliding window in a loop, the local features of all local areas are finally obtained ; where N num represents the total number of sliding times, and the calculation formula is:
[0065] .
[0066] Furthermore, in step three, the local area is serialized and the local cross-attention of the visible light and thermal infrared modalities is calculated. The specific steps are as follows:
[0067] Define local position encoding (Local Position Embedding) E ∈R 1×(ks×ks)×1 , broadcast it to the same dimension as S , and add it to all sliding windows to strengthen the local position information; through the local linear mapping layer, map the local area to obtain the local query Q local and the key value K local ; through the global linear mapping layer, map the original feature map to the global weight V global ;
[0068] Calculate the local cross-attention of the visible light and thermal infrared modalities using the following formula:
[0069] (a)
[0070] where d k represents the feature channel dimension.
[0071] Furthermore, to improve the calculation efficiency, select to calculate the weights of all blocks at once to obtain V∈R H×W×C , V represents the attention weight value, and through the same sliding window in VSwipe up to obtain for attention calculation.
[0072] Furthermore, to promote the fusion of different modalities, the visible light and thermal infrared modalities independently calculate the local cross-attention Attn. When calculating, V in formula (a) global comes from the other modality. That is, when calculating the local cross-attention of the visible light modality, use V of the thermal infrared modality global to calculate Attn. When calculating the local cross-attention of the thermal infrared modality, use V of the visible light modality global to calculate Attn;
[0073] Furthermore, in step three, the local cross-attention results of each sliding window are merged. The specific steps are as follows:
[0074] Divide the local attention result into N num local feature maps to obtain S' ∈ R ks×ks×C , S' denote the local feature maps, restore the local feature maps to the positions of the original feature maps corresponding to the sliding windows, calculate the average value of the overlapping parts of the local feature maps, add the local attention results according to the corresponding positions, and calculate the attention average value according to the coverage of the sliding windows in each block as the final local attention score.
[0075] Step four: Target tracking prediction;
[0076] The Transformer encoder at the final layer of the backbone network outputs the features of the visible light and thermal infrared modalities, which are merged using convolution and input to the prediction head for prediction to obtain the target position. Use common datasets for tracking tests to evaluate the tracking results.
[0077] Experimental example one: To better demonstrate the effectiveness of the target tracking method of the present invention, the method of the present invention and other RGBT target tracking methods are tested for tracking performance on the LasHeR, RGBT210, and RGBT234 datasets. The results are shown in Table 1.
[0078] Table 1 Tracking performance of the method of the present invention and other RGBT target tracking methods
[0079]
[0080] Among them, BAT (from the paper "Bi-directional Adapter for Multimodal Tracking", published in the AAAI 2024 international conference), CMD (from the paper "Efficient RGB-T Tracking via Cross-Modality Distillation", published in the CVPR 2023 international conference), CAT (from the paper "CAT: Challenge-Aware RGBT Tracking", published in the ECCV 2020 international conference), CAT++ (from the paper "RGBT Tracking via Challenge-Based Appearance Disentanglement and Interaction", published in the IEEE Transactions on Image Processing journal in 2024). DAPNet, DAFNet, MANet, mfDiMP, FANet, HMFT, APFNet, ProTrack, MACFT, ViPT, STMT, SDSTrack, OneTracker, M3PT are also trackers publicly disclosed in the prior art.
[0081] As can be seen from Table 1, the present invention exhibits excellent tracking performance on the three datasets of LasHeR, RGBT210, and RGBT234.
[0082] Experimental Example 2: Figure 5 This is the comparison result of the challenge attributes between the method of the present invention and other methods.
[0083] The challenge attributes include partial occlusion (PO), total occlusion (TO), hyaline occlusion (HO), motion blur (MB), low illumination (LI), high illumination (HI), abrupt illumination variation (AIV), low resolution (LR), deformation (DEF), background clutter (BC), similar appearance (SA), camera moving (CM), thermal crossover (TC), frame lost (FL), out of view (OV), fast motion (FM), scale variation (SV), aspect ratio change (ARC), and no occlusion (NO).
[0084] It can be seen that Figure 5 the method of the present invention performs excellently in most challenge attributes and only ranks second under the OV out-of-view challenge attribute. The method of the present invention far outperforms other trackers under the AIV abrupt illumination variation challenge attribute, demonstrating that the method of the present invention can more effectively utilize local features and cross-modal information to assist in decision-making.
[0085] Experimental Example 3: To verify the effectiveness of the convolutional attention fusion module of the present invention, the convolutional attention fusion module is replaced with a convolutional module, conventional attention, and windowed attention respectively, and compared with the baseline algorithm (OSTrack). The tracking performance detection results of different methods on the LasHeR dataset are shown in Table 2.
[0086] Table 2 Tracking performance detection results of different methods on the LasHeR dataset
[0087]
[0088] It can be seen from Table 2 that the above methods have all achieved an improvement in tracking performance, but the improvement brought by the convolutional attention fusion module is greater. Although conventional attention has a smaller number of parameters and computational complexity, the convolutional attention fusion module only sacrifices a small number of parameters and computational complexity to significantly improve the tracking performance, reflecting the superiority of the convolutional attention fusion module.
[0089] Experimental Example 4: To analyze the effects brought by different parameters of the convolutional attention fusion module, ablation experiments were conducted by modifying the module parameters. Among them, Experiment 1 was the baseline algorithm (OSTrack) for comparison. The results of the ablation experiments are shown in Table 3.
[0090] Table 3 Ablation Analysis of the Convolutional Attention Fusion Module on LasHeR
[0091]
[0092] Among them, in the padding experiment: 0 means no padding, 1 means padding 1 pixel point outward (up, down, left, and right), and the original 16×16 feature map becomes 18×18.
[0093] It can be seen from Table 3 that:
[0094] In Experiment 2, only window attention was used, and the impact on the tracker performance was extremely low.
[0095] In Experiment 3, with only local position encoding, there was a small improvement in performance, indicating that the convolutional attention fusion module can bring performance improvement even without introducing cross-modal information.
[0096] Experiment 4 mainly removed the local position encoding, and the performance decreased slightly compared to the complete model, indicating the effectiveness of the local position encoding in Step 3.
[0097] Experiment 5 was the method of the present invention, and the performance improvement was the most significant, indicating the importance of local features and cross-modal fusion for RGBT tracking.
[0098] In Experiment 6, padding was cancelled, that is, no zero-padding was performed on the window. In this case, the edge blocks of the two-dimensional features were only extracted once, and the utilization of local information was insufficient, so the performance decreased.
[0099] In Experiment 7, non-overlapping sliding windows were used to extract features, and due to the lack of interaction between windows, the performance also decreased.
[0100] Experimental Example 5: To analyze the influence of different window sizes on the tracking performance, the present invention conducted experiments by only modifying the window size under the condition of fixed padding and stride. The results are as Figure 8 shown.
[0101] Through Figure 8It can be seen that the present invention is not sensitive to the size of the sliding window, and the best performance is achieved when the window size is 2×2. When the window size increases, the performance first decreases and then increases. This is because it is difficult to extract discriminative local features and capture sufficient global features with a medium window size. Although the large window loses local information, it can better extract global features, so the performance is slightly better than that of the medium-sized window.
[0102] Experimental Example Six: The influence of the convolutional attention fusion module at different layers on the tracking performance is shown in Table 4.
[0103] Table 4 Influence of the layer where the convolutional attention fusion module is located on the tracking performance
[0104]
[0105] It can be seen from Table 4 that when the convolutional attention fusion module is embedded in layers 4, 7, and 10, the tracking performance reaches the highest. However, when the convolutional attention fusion module is embedded in all layers, the performance decreases instead. This is because the convolutional attention fusion module focuses on the extraction and fusion of local features, and overusing the convolutional attention fusion module will cause local features to dominate, resulting in the model ignoring more discriminative global features.
[0106] Experimental Example Seven: To verify the generality of the convolutional attention module, Table 5 shows its usage effect in single-modal tracking. Using Figure 3 the shown convolutional self-attention, OSTrack is selected as the baseline algorithm, its built-in candidate elimination module is deleted, and the convolutional attention fusion module is embedded in layers 4, 7, and 10 of OSTrack with the same parameters.
[0107] Table 5 Effect of the convolutional attention module on the single-modal tracking model
[0108]
[0109] LaSOT (Large-scale Single Object Tracking dataset, from the paper LaSOT: A High-quality Benchmark for Large-scale Single Object Tracking, published in the CVPR 2019 international conference).
[0110] It can be seen from Table 5 that the convolutional attention module still has a certain performance improvement for single-modal tracking, indicating that the convolutional attention module proposed in the present invention has a certain universality for object tracking.
[0111] Experimental Example Eight:
[0112] Figure 6This is the comparison result of the accuracy of the present invention with other methods on the LasHeR dataset. It can be seen that the present invention achieves the minimum error at most thresholds, indicating that the present invention has a high accuracy in predicting the target position.
[0113] Figure 7 This is the comparison result of the success rate of the present invention with other methods on the LasHeR dataset. It can be seen that the present invention has a high overlap rate at most thresholds, indicating that the present invention has a good success rate in predicting the target size.
[0114] Figure 9 - Figure 11 This is the visual tracking result of the method of the present invention and other RGB-T target tracking methods. It can be seen that the present invention can accurately track objects in challenging scenarios such as small targets and similar targets, and has good performance. Among them, TBSI (from the paper "Bridging Search Region Interaction with Template for RGB-T Tracking", published in the CVPR 2023 international conference).
[0115] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for RGBT object tracking based on convolutional attention fusion, characterized in that, It includes the following steps: Step 1: Video preprocessing; Randomly select a video sequence from the dataset, where each frame is an image; select the rectangular area where the target is located at the same position in the first frame of the visible light and thermal infrared modalities, scale this area, and save it as the target template of the video sequence; starting from the second frame, take the target position of the previous frame as the center point, select a square area with a range larger than the target area, and scale it as the search area of the current frame; Step 2: Feature extraction; Cut the target template and search area of the visible light and thermal infrared modalities into several blocks and unfold and splice them. Map the image information into a one-dimensional feature sequence through a linear mapping layer, and add global position encoding to it; use the Transformer encoder of the backbone network with shared parameters to extract features for the features of the two modalities respectively; The backbone network uses OSTrack as the baseline model and is extended to a two-branch structure. Let the original input visible light and thermal infrared features be X r and X t . The i -th layer Transformer encoder is denoted as Encoder i . The feature extraction process can be formulated as: ; Among them, represents the visible light feature of the i-th layer, represents the thermal infrared feature of the i-th layer, represents the visible light feature of the (i + 1)-th layer, represents the thermal infrared feature of the (i + 1)-th layer; Step 3: Local feature enhancement and fusion; Insert a convolutional attention fusion module between the Transformer encoders: perform two-dimensional processing on the one-dimensional feature sequence through a sliding window, then perform zero-padding around it, use the sliding window to slide cyclically starting from the upper left corner, and the sliding window traverses the entire feature map to select the local area; serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities; finally, merge the local cross-attention results of each sliding window; Select the local area through a sliding window. The specific steps are as follows: For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , expand it into two-dimensional features and extend it to ( H +2× padding )×( W +2× padding ) by zero-padding, where H and W, C are the height, width and number of channels of the original feature map respectively, padding is the padding size; divide the extended area into p×p blocks, and regard each block as a feature point with a dimension of P ∈ R p×p×C , where P refers to any feature point; then use a sliding window with a side length of ks to slide from the upper left corner of the extended feature map, with a step size of stride , and select the local area X left:left+ks,top:top+ks , where ( left , top ) is the upper left corner coordinate, and the sliding window is represented as X' ∈ R ks×ks×C ; the sliding window slides continuously in a loop, and finally obtains the local features of all local areas; among them, N num represents the total number of sliding times, and the calculation formula is: ; Serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities. The specific steps are as follows: Define local positional encoding E ∈R 1×(ks×ks)×1 , broadcast it to the same dimension as S , and add it to all sliding windows; through the local linear mapping layer, map the local area to obtain the local query Q local and key-value K local ; through the global linear mapping layer, map the original feature map to the global weight V global ; Calculate the local cross-attention of the visible light and thermal infrared modalities using the following formula: (a) Among them, d k represents the dimension of the feature channel; Merge the local cross-attention results of each sliding window. The specific steps are as follows: Divide the local attention result into N num local feature maps, obtaining S' ∈ R ks×ks×C , S' where represents the local feature maps. Restore the local feature maps to the positions of the original feature maps corresponding to the sliding window, calculate the average value of the overlapping parts of the local feature maps, add the local attention results according to the corresponding positions, and calculate the attention average value based on the coverage of the sliding window in each block as the final local attention score; Step 4: Target tracking prediction: The Transformer encoder of the final layer of the backbone network outputs the features of the visible light and thermal infrared modalities, uses convolution to merge them, and inputs them into the prediction head for prediction to obtain the target position.
2. The RGBT object tracking method based on convolutional attention fusion according to claim 1, wherein In Step 1, the LasHeR dataset is used as the dataset.
3. The RGBT object tracking method based on convolutional attention fusion according to claim 1, wherein Insert a convolutional attention fusion module between the 4th, 7th, and 10th layers of the Transformer encoder of the backbone network.
4. The RGBT object tracking method based on convolutional attention fusion according to claim 1, wherein, Obtained by the method of calculating the weights of all blocks at once V∈R H×W×C , V denotes the attention weight value, which slides on V by the same sliding window to obtain .
5. The RGBT object tracking method based on convolutional attention fusion according to claim 4, characterized in that, The local cross-attention Attn is calculated independently for the visible light and thermal infrared modalities. When calculating the local cross-attention of the visible light modality, the V of the thermal infrared modality is used to calculate Attn. global When calculating the local cross-attention of the thermal infrared modality, the V of the visible light modality is used global to calculate Attn.
Citation Information
Patent Citations
Target tracking method combining feature enhancement and template updating
CN115205730A
Feature fusion target tracking method based on attention mechanism
CN116310683A