RGBT target tracking method based on convolution attention fusion
By introducing a convolutional attention fusion module into the RGBT target tracking method, the problem of insufficient utilization of local features and cross-modal features in the prior art is solved, and a more stable and efficient RGBT target tracking performance is achieved.
Patent Information
- Application Number
- CN202510582564.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing RGBT target tracking method focuses on modal global fusion, and fails to fully utilize local features and exert the potential of convolution and attention mechanisms, resulting in insufficient robustness and accuracy.
A RGBT target tracking method based on convolutional attention fusion is designed to improve the use of local features by the tracker through the convolutional attention fusion module, and promote cross-modal feature fusion to maintain attention to global information.
Through local feature enhancement and cross-modal interaction, the performance of RGBT target tracking is improved, the model's utilization of visible light and thermal infrared modal features is enhanced, and the impact of single mode defects on tracking performance is avoided.
Smart Images

Figure CN120088292A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly relates to an RGBT object tracking method based on convolutional attention fusion. Background Art
[0002] Object tracking is one of the important research contents in the field of computer vision, aiming to continuously track an object in a given video sequence, and it has been widely used in many fields such as video surveillance, human-computer interaction, and visual navigation. Early tracking algorithms tended to use handcrafted features to track objects based on domain knowledge and experience. However, the performance of these methods is difficult to meet the requirements in real scenarios. Since the proposal of deep learning, with its powerful feature modeling ability, it has gradually replaced traditional algorithms and become the mainstream method in the field of object tracking.
[0003] In recent years, with the proposal of deep learning models such as convolutional neural networks and Transformers, significant progress has been made in object tracking methods based on deep learning. These methods learn image features in a data-driven manner, improving the performance of object tracking. RGB object tracking only uses the visible light modality to track objects and still faces many challenges. Visible light-based tracking methods are vulnerable to factors such as light intensity, object color, and weather conditions. Some researchers have tried to introduce data from other modalities and fuse it with visible light data to improve the accuracy of the tracking algorithm using the complementary information of different modalities. Infrared light has a thermal effect, can reflect the temperature of an object, and is not affected by light, enabling the identification of objects at night. Therefore, infrared light has broad application prospects in the field of object tracking. RGBT object tracking comprehensively uses thermal infrared information and visible light information to perform object tracking. However, thermal infrared images have defects such as severe noise, low resolution, unclear objects, inconspicuous texture features, and difficulty in penetrating transparent objects. In contrast, although visible light provides rich texture information and color information, it is easily affected by lighting conditions.
[0004] In the field of RGBT object tracking, how to effectively fuse the information of the visible light and thermal infrared modalities and utilize their respective advantages to improve the robustness and accuracy of tracking is a particularly important research point. Existing RGBT object tracking methods can be divided into pure convolutional-based methods, convolutional-Transformer hybrid methods, and pure Transformer-based methods. However, most methods still focus on the global fusion of modalities and fail to consider the fusion and utilization of local features, nor fully exploit the potential of convolution and the attention mechanism.
[0005] Based on this, the present invention designs an RGBT object tracking method based on convolutional attention fusion to solve the above problems. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the present invention provides an RGBT object tracking method based on convolutional attention fusion. By designing a convolutional attention fusion module, the tracker's utilization of local features is improved, cross-modal feature fusion is promoted, and attention to global information is maintained, thereby achieving more stable RGBT object tracking.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: An RGBT object tracking method based on convolutional attention fusion, comprising the following steps: Step 1, video preprocessing; Randomly select a video sequence from the dataset, where each frame is an image; select the rectangular area where the target is located at the same position in the first frame of the visible light and thermal infrared modalities, scale this area, and save it as the target template of the video sequence; starting from the second frame, use the target position of the previous frame as the center point, select a square area with a range larger than the target area, and scale it as the search area of the current frame; Step 2, feature extraction; Cut the target template and search area of the visible light and thermal infrared modalities into several blocks and unfold and splice them. Map the image information into a one-dimensional feature sequence through a linear mapping layer, and add global position encoding to it; use the Transformer encoder of the backbone network with shared parameters to extract features for the features of the two modalities respectively; Step 3, local feature enhancement and fusion; Insert a convolutional attention fusion module between the Transformer encoders: perform two-dimensional processing on the one-dimensional feature sequence through a sliding window, then perform zero-padding around it, use the sliding window to slide cyclically starting from the upper left corner, the sliding window traverses the entire feature map, and select the local area; serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities; finally, merge the local cross-attention results of each sliding window; Step 4, object tracking prediction.
[0008] Furthermore, in step 1, the LasHeR dataset is used as the dataset.
[0009] Furthermore, the backbone network in step 2 uses OSTrack as the baseline model and is extended to a two-branch, and the original input visible light and thermal infrared features are respectively X r 、 X t ,the i layer Transformer encoder is denoted as Encoder i, the feature extraction process can be formulated as follows: ; where, represents the visible light feature of the i-th layer, represents the thermal infrared feature of the i-th layer, represents the visible light feature of the (i + 1)-th layer, represents the thermal infrared feature of the (i + 1)-th layer.
[0010] Furthermore, convolutional attention fusion modules are inserted into the 4th, 7th, and 10th layers between the Transformer encoders of the backbone network.
[0011] Furthermore, in step three, a local area is selected through a sliding window, and the specific steps are as follows: For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , it is expanded into a two-dimensional feature and extended to ( H +2× padding )×( W +2× padding ) by zero padding, where H , W, C are the height, width, and number of channels of the original feature map respectively, padding is the padding size; the expanded area is divided into p×p blocks, and each block is regarded as a feature point with a dimension of P ∈ R p×p×C , where P refers to any feature point; then a sliding window with a side length of ks starts sliding from the upper left corner of the expanded feature map with a step size of stride , and a local area X left:left+ks,top:top+ks is selected, where ([[]] left , top ) is the upper left corner coordinate, and the sliding window is represented as X' ∈ R ks×ks×C ; the sliding window slides continuously in a loop, and finally the local features of all local areas are obtained; where N num represents the total number of sliding times, and the calculation formula is: .
[0012] Furthermore, in step three, the local areas are serialized and the local cross-attention of the visible light and thermal infrared modalities is calculated, and the specific steps are as follows: Define the local position encodingE ∈R 1×(ks×ks)×1 , broadcast it to the same dimension as S and add it to all sliding windows; through the local linear mapping layer, map the local area to obtain the local query Q local and the key-value K local ; through the global linear mapping layer, map the original feature map to the global weight V global ; Calculate the local cross-attention of the visible light and thermal infrared modalities, using the following formula: (a) where d k represents the feature channel dimension.
[0013] Further, use the method of calculating the weights of all blocks at once to obtain V ∈ R H×W×C , V represents the attention weight value, slide on V through the same sliding window to obtain .
[0014] Further, the visible light and thermal infrared modalities independently calculate the local cross-attention Attn. When calculating the local cross-attention of the visible light modality, use the V of the thermal infrared modality global to calculate Attn. When calculating the local cross-attention of the thermal infrared modality, use the V of the visible light modality global to calculate Attn.
[0015] Further, in step three, merge the local cross-attention results of each sliding window. The specific steps are as follows: Divide the local attention result into N num local feature maps to obtain S' ∈ R ks ×ks×C , S' represents the local feature map, restore the local feature map to the position of the original feature map corresponding to the sliding window, take the average value of the overlapping parts of the local feature maps, add the local attention results according to the corresponding positions, and calculate the attention average value according to the coverage of the sliding window in each block as the final local attention score.
[0016] Furthermore, Step 4 is specifically as follows: The Transformer encoder at the final layer of the backbone network outputs the features of the visible light and thermal infrared modalities, which are merged using convolution and input into the prediction head for prediction to obtain the target position.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on the Transformer tracking architecture, the present invention integrates the ideas of convolution and attention, designs an efficient convolutional attention fusion module, thereby completing local feature enhancement and cross-modal interaction; through complementarity with the global features of the ViT network, the RGBT target tracking performance is effectively improved.
[0018] The present invention enhances the robustness of the target tracking task through multi-modal fusion. The model extracts features from both the visible light and thermal infrared modalities simultaneously, and through the convolutional attention fusion module, realizes local feature enhancement and cross-modal feature interaction, avoiding the influence of single-modal defects on the tracking performance.
[0019] The present invention can be used for any RGBT video sequence. The proposed convolutional attention fusion module can adjust padding, stride, window size, etc., and has strong adjustability. A better combination can be selected according to specific situations. In addition, the convolutional attention fusion module can be applied to most ViT-based trackers.
[0020] In the present invention, video sequences are randomly selected from the LasHeR dataset, and the target templates are calibrated; the bimodal images are scaled to a unified size and cut into multiple blocks, and the features of each block are encoded using linear mapping and flattened into a one-dimensional feature sequence; the Transformer encoder is iteratively used to extract the sequence features; during the feature extraction process, the convolutional attention fusion module is interspersed. First, the one-dimensional feature sequence is merged into a two-dimensional feature map, local regions are selected on the feature map through a sliding window, and the local features and cross-modal features are enhanced through the attention mechanism; the enhanced multi-modal features are merged through convolution and input into the prediction head for tracking prediction; the performance of the present invention on multiple video tracking datasets is better than other advanced RGBT tracking algorithms.
[0021] The present invention can realize more effective utilization of local features and cross-modal features, and solves the deficiencies of existing RGBT trackers in the utilization of local features and cross-modal interaction. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a comparison chart of conventional attention, windowed attention, and convolutional self-attention. Among them: (a) is conventional attention, which directly calculates attention on the complete feature map. (b) is windowed attention, which divides the feature map into multiple non-overlapping windows, calculates local attention within the windows, and merges adjacent windows as the number of layers deepens. (c) is convolutional self-attention, which calculates local attention of adjacent regions through overlapping sliding windows.
[0024] Figure 2 It is the overall structure diagram of the RGBT object tracking model of the present invention.
[0025] Figure 3 It is a schematic diagram of convolutional self-attention.
[0026] Figure 4 It is a schematic diagram of convolutional cross-attention.
[0027] Figure 5 It is the comparison result of the challenge attributes between the method of the present invention and other methods.
[0028] Figure 6 It is the comparison result of the accuracy between the present invention on the LasHeR dataset and other methods.
[0029] Figure 7 It is the comparison result of the success rate between the present invention on the LasHeR dataset and other methods.
[0030] Figure 8 It is the influence of different window sizes on the tracking performance.
[0031] Figure 9 It is the first example of the visual tracking results of the method of the present invention and other RGBT object tracking methods (the fourth boy in the first column).
[0032] Figure 10 It is the second example of the visual tracking results of the method of the present invention and other RGBT object tracking methods (shooting at the basket three times).
[0033] Figure 11 It is the third example of the visual tracking results of the method of the present invention and other RGBT object tracking methods (the boy playing with the mobile phone). Detailed implementation manners
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment 1: An RGBT object tracking method based on convolutional attention fusion, comprising the following steps: Step 1: Video preprocessing; In a common RGBT object tracking dataset, a video sequence is randomly selected, where each frame is an image; in the same position of the first frame in the visible light and thermal infrared modalities, a rectangular area where the object is located is selected, and this area is scaled to 128×128 pixels and saved as the object template for this video sequence; starting from the second frame, using the object position of the previous frame as the center point, assuming the width and height of the area where the object is located in the previous frame are w and h respectively, a square area with a side length of is selected and scaled to 256×256 pixels as the search area for the current frame; Further, in Step 1, the video sequence in the LasHeR dataset is selected as the tracking video; the LasHeR dataset contains a total of 1224 video sequences and 734.8K frames, with rich data resources and wide coverage of real-world scenarios, which are sufficient for model training.
[0036] Step 2: Feature extraction; The object template and search area in the visible light and thermal infrared modalities are cut into several 16×16 blocks and unfolded and spliced, and the image information is mapped into a one-dimensional feature sequence through a linear mapping layer, and global position encoding is added to it; the Transformer encoders of the backbone network with shared parameters are used to extract features for the features of the two modalities respectively.
[0037] The size of the search area is 256×256, which is cut into blocks of size 16×16 in Step 2, a total of 16×16 blocks. After passing through the linear mapping layer, the block information is mapped to a point to obtain a feature map. Therefore, the size of the feature map is 16×16, and each feature point in it corresponds to the original 16×16 block.
[0038] Further, in Step 2, the backbone network uses OSTrack as the baseline model and is extended to a dual-branch, and the original input visible light and thermal infrared features are respectively X r 、 X t , the iThe layer Transformer encoder is denoted as Encoder i , and the feature extraction process can be formulated as: ; where represents the visible light feature of the i-th layer, represents the thermal infrared feature of the i-th layer, represents the visible light feature of the (i + 1)-th layer, represents the thermal infrared feature of the (i + 1)-th layer.
[0039] Step 3: Local feature enhancement and fusion; Insert a convolutional attention fusion module at specified positions (the 4th, 7th, and 10th layers) between the Transformer encoders of the backbone network. The steps are as follows: Two-dimensionally process the one-dimensional feature sequence through a sliding window, then perform zero-padding around it, use the sliding window to slide cyclically starting from the upper left corner, and the sliding window traverses the entire feature map to select a local area; Serialize the local area and calculate the local cross-attention of the visible light and thermal infrared modalities; Finally, merge the local cross-attention results of each sliding window to enhance local information and cross-modal feature interaction.
[0040] The convolutional attention fusion module is implemented using convolutional cross-attention. As a plug-and-play module, the convolutional attention fusion module can be embedded into most trackers using ViT and is applicable to both single-modal and dual-modal scenarios. The inventive method can achieve relatively stable tracking for any RGBT video sequence and has high performance in most challenging scenarios.
[0041] Furthermore, in Step 3, the local area is selected through a sliding window, and the specific steps are as follows: For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , expand it into a two-dimensional feature and expand it to ([[]]END]] H +2× padding )×([[]]END]] W +2× padding ) through zero-padding, where H , W, C are the height, width, and number of channels of the original feature map respectively, padding is the padding size; Divide the expanded area into p×p blocks, and regard each block as a feature point with a dimension of P ∈ R p×p×C , where P refers to any feature point; Subsequently, use a side length of ksThe sliding window (e.g., a sliding window of size 2×2) slides from the upper left corner of the expanded feature map with a stride of stride (e.g., 1), and selects the local region X left:left+ks,top:top+ks , where ( left , top ) are the upper left coordinates, and the sliding window is represented as X' ∈ R ks×ks×C ; By continuously sliding the sliding window in a loop, the local features of all local regions are finally obtained ; Among them, N num represents the total number of sliding times, and the calculation formula is: .
[0042] Further, in step three, serialize the local regions and calculate the local cross-attention of the visible light and thermal infrared modalities. The specific steps are as follows: Define the local position encoding (Local Position Embedding) E ∈R 1×(ks×ks)×1 , broadcast it to the same dimension as S , and add it to all the sliding windows to strengthen the local position information; through the local linear mapping layer, map the local region to obtain the local query Q local and the key value K local ; Through the global linear mapping layer, map the original feature map to the global weight V global ; Calculate the local cross-attention of the visible light and thermal infrared modalities, using the following formula: (a) where d k represents the feature channel dimension.
[0043] Further, in order to improve the calculation efficiency, select to calculate the weights of all blocks at one time to obtain V ∈ R H×W×C , V represents the attention weight value, and slides on V through the same sliding window to obtain for attention calculation.
[0044] Further, in order to promote the fusion of different modalities, the visible light and thermal infrared modalities independently calculate the local cross-attention Attn. When calculating, V in formula (a) globalFrom the other modality, that is, when calculating the local cross-attention of the visible light modality, use the V of the thermal infrared modality global to calculate Attn. When calculating the local cross-attention of the thermal infrared modality, use the V of the visible light modality global to calculate Attn; Furthermore, in step three, merge the local cross-attention results of each sliding window. The specific steps are as follows: The local attention result is divided into N num local feature maps to obtain S' ∈ R ks×ks×C , S' denotes the local feature map, and restore the local feature map to the position of the original feature map corresponding to the sliding window. Take the average value of the overlapping parts of the local feature maps, add the local attention results according to the corresponding positions, and calculate the attention average value according to the coverage of the sliding window in each block as the final local attention score.
[0045] Step four, target tracking prediction; The Transformer encoder of the final layer of the backbone network outputs the features of the visible light and thermal infrared modalities, which are merged by convolution and input into the prediction head for prediction to obtain the target position. Use common datasets for tracking tests to evaluate the tracking results.
[0046] Experimental example one: In order to better reflect the effectiveness of the target tracking method of the present invention, detect the tracking performance of the method of the present invention and other RGBT target tracking methods on the LasHeR, RGBT210, and RGBT234 datasets. The results are shown in Table 1.
[0047] Table 1 Tracking performance of the method of the present invention and other RGBT target tracking methods
[0048] Among them, BAT (from the paper "Bi-directional Adapter for Multimodal Tracking" published in the AAAI 2024 international conference), CMD (from the paper "Efficient RGB-T Tracking via Cross-Modality Distillation" published in the CVPR 2023 international conference), CAT (from the paper "CAT: Challenge-Aware RGBT Tracking" published in the ECCV 2020 international conference), CAT++ (from the paper "RGBT Tracking via Challenge-Based Appearance Disentanglement and Interaction" published in the IEEE Transactions on Image Processing journal in 2024). DAPNet, DAFNet, MANet, mfDiMP, FANet, HMFT, APFNet, ProTrack, MACFT, ViPT, STMT, SDSTrack, OneTracker, and M3PT are also trackers already disclosed in the prior art.
[0049] As can be seen from Table 1, the present invention exhibits excellent tracking performance on the three datasets of LasHeR, RGBT210, and RGBT234.
[0050] Experimental Example 2: Figure 5 This is the comparison result of the challenge attributes between the method of the present invention and other methods.
[0051] The challenge attributes include partial occlusion (PO), total occlusion (TO), hyaline occlusion (HO), motion blur (MB), low illumination (LI), high illumination (HI), abrupt illumination variation (AIV), low resolution (LR), deformation (DEF), background clutter (BC), similar appearance (SA), camera moving (CM), thermal crossover (TC), frame lost (FL), out of view (OV), fast motion (FM), scale variation (SV), aspect ratio change (ARC), and no occlusion (NO).
[0052] Through Figure 5 It can be seen that the method of the present invention performs excellently in most challenge attributes and only ranks second under the OV out-of-view challenge attribute. The method of the present invention far outperforms other trackers under the AIV abrupt illumination variation challenge attribute, demonstrating that the method of the present invention can more effectively utilize local features and cross-modal information to assist in decision-making.
[0053] Experimental Example 3: To verify the effectiveness of the convolutional attention fusion module of the present invention, the convolutional attention fusion module was replaced with a convolutional module, conventional attention, and windowed attention respectively, and compared with the baseline algorithm (OSTrack). The tracking performance detection results of different methods on the LasHeR dataset are shown in Table 2.
[0054] Table 2 Tracking performance detection results of different methods on the LasHeR dataset
[0055] It can be seen from Table 2 that the above methods have all achieved an improvement in tracking performance, but the improvement brought by the convolutional attention fusion module is greater. Although conventional attention has a smaller number of parameters and computational complexity, the convolutional attention fusion module only sacrifices a small number of parameters and computational complexity to significantly improve the tracking performance, reflecting the superiority of the convolutional attention fusion module.
[0056] Experimental Example 4: To analyze the impact of different parameters of the convolutional attention fusion module, ablation experiments were conducted by modifying the module parameters. Among them, Experiment 1 was the baseline algorithm (OSTrack) for comparison. The results of the ablation experiments are shown in Table 3.
[0057] Table 3 Ablation Analysis of the Convolutional Attention Fusion Module on LasHeR
[0058] Among them, in the padding experiment: 0 means no padding, 1 means padding 1 pixel point outward (up, down, left, and right), and the original 16×16 feature map becomes 18×18.
[0059] It can be seen from Table 3 that: In Experiment 2, only window attention was used, and the impact on the tracker performance was extremely low.
[0060] In Experiment 3, with only local position encoding, there was a small improvement in performance, indicating that the convolutional attention fusion module can bring performance improvement even without introducing cross-modal information.
[0061] Experiment 4 mainly removed the local position encoding, and the performance decreased slightly compared to the complete model, indicating the effectiveness of the local position encoding in Step 3.
[0062] Experiment 5 was the method of the present invention, and the performance improvement was the most significant, indicating the importance of local features and cross-modal fusion for RGBT tracking.
[0063] In Experiment 6, padding was cancelled, that is, no zero-padding was performed on the window. In this case, the edge blocks of the two-dimensional features were only extracted once, and the local information was not fully utilized, so the performance decreased.
[0064] In Experiment 7, non-overlapping sliding windows were used to extract features. Due to the lack of interaction between windows, the performance also decreased.
[0065] Experimental Example 5: To analyze the impact of different window sizes on the tracking performance, the present invention conducted experiments by only modifying the window size under the condition of fixed padding and stride. The results are as Figure 8 shown.
[0066] Through Figure 8 it can be seen that the present invention is not sensitive to the size of the sliding window, and the best performance is achieved when the window size is 2×2. When the window size increases, the performance first decreases and then increases. This is because medium window sizes are difficult to extract discriminative local features and capture sufficient global features. While large windows lose local information, they can better extract global features, so the performance is slightly better than that of medium-sized windows.
[0067] Experimental Example Six: The influence of the convolutional attention fusion module at different layers on the tracking performance is shown in Table 4.
[0068] Table 4 Influence of the layer where the convolutional attention fusion module is located on the tracking performance
[0069] It can be seen from Table 4 that when the convolutional attention fusion module is embedded in layers 4, 7, and 10, the tracking performance reaches the highest. However, when the convolutional attention fusion module is embedded in all layers, the performance decreases instead. This is because the convolutional attention fusion module focuses on the extraction and fusion of local features, and overusing the convolutional attention fusion module will cause local features to dominate, resulting in the model ignoring more discriminative global features.
[0070] Experimental Example Seven: To verify the generality of the convolutional attention module, Table 5 shows its usage effect in single-modal tracking. The convolutional self-attention shown in Figure 3 is adopted. Select OSTrack as the baseline algorithm, delete its built-in candidate elimination module, and embed the convolutional attention fusion module into layers 4, 7, and 10 of OSTrack with the same parameters.
[0071] Table 5 Effect of the convolutional attention module on the single-modal tracking model
[0072] LaSOT (Large-scale Single Object Tracking dataset, from the paper LaSOT: A High-quality Benchmark for Large-scale Single Object Tracking, published in the CVPR 2019 international conference).
[0073] It can be seen from Table 5 that the convolutional attention module still has a certain performance improvement in single-modal tracking, indicating that the convolutional attention module proposed in the present invention has a certain universality for target tracking.
[0074] Experimental Example Eight: Figure 6 This is the accuracy comparison result of the present invention with other methods on the LasHeR dataset. It can be seen that the present invention achieves the minimum error at most thresholds, indicating that the present invention has a high accuracy in predicting the target position.
[0075] Figure 7 This is the success rate comparison result of the present invention with other methods on the LasHeR dataset. It can be seen that the present invention has a high overlap rate at most thresholds, indicating that the present invention has a good success rate in predicting the target size.
[0076] Figure 9 - Figure 11 The figure shows the visual tracking results of the method of the present invention and other RGB-T object tracking methods. It can be seen that the present invention can accurately track objects in challenging scenarios such as small objects and similar objects, and has good performance. Among them, TBSI (from the paper "Bridging Search Region Interaction with Template for RGB-T Tracking", published in the CVPR 2023 international conference).
[0077] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements will not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A RGBT target tracking method based on convolutional attention fusion, characterized in that: The following steps are involved: Step 1: Video preprocessing; A video sequence is randomly selected from the data set, where each frame is an image; a rectangular area where the target is located is selected at the same position in the first frame of the visible light and thermal infrared modes, the area is scaled, and saved as the target template of the video sequence; starting from the second frame, a square area larger than the target area is selected with the target position of the previous frame as the center point, and the area is scaled as the search area of the current frame; Step 2: Feature extraction; The target template and search area of the visible light and thermal infrared modalities are cut into several blocks and then expanded and spliced. The image information is mapped into a one-dimensional feature sequence through a linear mapping layer, and a global position encoding is added to it. The Transformer encoder of the backbone network with shared parameters is used to extract features from the features of the two modalities. Step 3: Local feature enhancement and fusion; Insert a convolutional attention fusion module between the Transformer encoders: convert the one-dimensional feature sequence into two-dimensional one through a sliding window, then fill it with zeros around it, and use the sliding window to slide cyclically from the upper left corner. The sliding window traverses the entire feature map and selects a local area. Serialize the local area and calculate the local cross attention of visible light and thermal infrared modalities; finally, merge the local cross attention results of each sliding window; Step 4: Target tracking prediction.
2. The RGBT target tracking method based on convolutional attention fusion according to claim 1, characterized in that: In step 1, the data set uses the LasHeR data set.
3. The RGBT target tracking method based on convolutional attention fusion according to claim 1, characterized in that: The backbone network in step 2 uses OSTrack as the baseline model and expands it into a dual-branch model, making the original input visible light and thermal infrared features respectively X r , X t , No. i The layer Transformer encoder is represented as Encoder i , the feature extraction process can be formulated as: ; in, represents the i-th layer visible light feature, represents the thermal infrared characteristics of the i-th layer, represents the (i+1)th layer of visible light features, Represents the thermal infrared characteristics of the (i+1)th layer.
4. The RGBT target tracking method based on convolutional attention fusion according to claim 1, characterized in that: The convolutional attention fusion module is inserted into the 4th, 7th, and 10th layers between the Transformer encoders of the backbone network.
5. The RGBT target tracking method based on convolutional attention fusion according to claim 1, characterized in that: In step 3, a local area is selected by sliding the window. The specific steps are as follows: For the input one-dimensional linear feature sequence X ∈ R (H×W)×C , expand it into a two-dimensional feature, and expand it to ( H +2× padding )×( W +2× padding ), in, H , W.C. are the height, width and number of channels of the original feature map, respectively. padding is the filling size; the expanded area is divided into p×p blocks, each block is considered as a P ∈ R p×p×C The characteristic points of P Refers to any feature point; then use the side length ks The sliding window starts sliding from the upper left corner of the expanded feature map, with a step size of stride , select a local area X left:left+ks,top:top+ks ,in( left , top ) is the coordinate of the upper left corner, and the sliding window is expressed as X' ∈ R ks×ks×C ; The sliding window continuously slides in a cycle, and finally obtains the local features of all local areas ;in, N num Represents the total number of slides, calculated as: 。 6. The RGBT target tracking method based on convolutional attention fusion according to claim 5, characterized in that: In step 3, the local regions are serialized and the local cross attention of the visible light and thermal infrared modalities is calculated. The specific steps are as follows: Defining local position encoding E ∈R 1×(ks×ks)×1 , broadcast it to S The same dimension is added to all sliding windows; Through the local linear mapping layer, the local area is mapped to obtain the local query Q local With key value K local ; Through the global linear mapping layer, the original feature map is mapped to the global weight V global ; The local cross attention between visible light and thermal infrared modalities is calculated using the following formula: (a) Among them, d k Represents the feature channel dimension.
7. The RGBT target tracking method based on convolutional attention fusion according to claim 6, characterized in that: The method of calculating the weight of all blocks at once is used to obtain V∈R H×W×C , V Represents the attention weight value, through the same sliding window in V Swipe up to get .
8. The RGBT target tracking method based on convolutional attention fusion according to claim 7, characterized in that: The local cross attention Attn is calculated independently for the visible light and thermal infrared modalities. When calculating the local cross attention of the visible light modality, the V of the thermal infrared modality is used global To calculate Attn, we use V of the visible light modality when calculating the local cross attention of the thermal infrared modality. global To calculate Attn.
9. The RGBT target tracking method based on convolutional attention fusion according to claim 8, characterized in that: In step 3, the local cross-attention results of each sliding window are merged. The specific steps are as follows: The local attention result Split into N num local feature maps, and we get S' ∈ R ks×ks×C , S' Represent the local feature map, and restore the local feature map to the position of the original feature map corresponding to the sliding window, average the overlapping parts of the local feature map, add the local attention results according to the corresponding positions, and calculate the attention average value according to the coverage of the sliding window in each block as the final local attention score.
10. The RGBT target tracking method based on convolutional attention fusion according to any one of claims 1 to 9, characterized in that: Step 4 is as follows: the Transformer encoder of the final layer of the backbone network outputs two modal features, visible light and thermal infrared, which are combined using convolution and input into the prediction head for prediction to obtain the target position.
Citation Information
Patent Citations
Target tracking method combining feature enhancement and template updating
CN115205730A
Feature fusion target tracking method based on attention mechanism
CN116310683A
RGBT real-time target tracking method and device based on feature enhancement fusion
CN117593334A
RGBT target tracking method based on modal perception feature learning
CN119068016A
Target tracking in medical image data
US20240296935A1
Cited By
Track platform target tracking method and device based on visible light-infrared image
CN120279065A
Track station target tracking method and device based on visible light-infrared image
CN120279065B
Highway engineering intelligent traffic flow prediction method based on big data
CN120496325A
Intelligent Traffic Flow Prediction Method for Highway Engineering Based on Big Data
CN120496325B