RGBT tracking method based on target adaptive text guide visual fusion
By introducing target adaptive text description and multimodal feature fusion technology into RGB-T tracker, the problem of poor performance of pure visual trackers in complex environments is solved, and higher tracking accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510510767.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing pure vision-based RGB-T trackers are not effective when facing challenges such as camera motion, motion blur, scale changes, target appearance changes, occlusion and background interference.
Using a visual fusion method based on target adaptive text guidance, a target text description is generated through the BLIP model, and combined with a visual encoder and a text encoder, a target text adaptive enhancement module with a sparse attention mechanism and a multimodal sharing and complementary information prompter are designed to enhance the expression ability of multimodal features.
It improves tracking accuracy and robustness, effectively responds to challenges such as camera motion, motion blur, and scale changes, and improves the performance of the tracker in complex environments.
Smart Images

Figure CN120031918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to an RGBT tracking method based on target adaptive text-guided visual fusion. Background Art
[0002] As thermal infrared devices are gradually deployed in smart devices and target tracking technology is widely used in robotics, video surveillance, autonomous driving and other fields, target tracking based on visible light and thermal infrared images (RGB-T tracking) has gradually become a highly anticipated research direction. Researchers are committed to constructing a robust tracking algorithm to cope with challenges that may arise during the tracking process, such as occlusion, camera motion, illumination changes, and scale changes.
[0003] In the past, most deep RGB-T trackers have focused on integrating information from both modalities to achieve efficient tracking performance. Some methods jointly learn the shared and specific information of visible light and thermal infrared modalities to improve the representation ability of multimodal features, and then perform feature fusion to achieve robust tracking. Other methods build attribute-aware networks, use a general branch or multiple independent branches to learn multimodal feature representations under specific attributes, and perform adaptive fusion of these attribute features to cope with challenges in various scenarios, thereby achieving robust tracking.
[0004] Recently popular high-performance RGB-T trackers are mainly divided into two categories. The first type of tracker uses the attention mechanism of the Transformer architecture to construct a multimodal fusion module, which allows cross-modal interaction between RGB and TIR, and fuses single-modal information into multimodal representation, thereby achieving robust tracking; the second type of tracker achieves efficient tracking performance through a multimodal tracking framework based on visual cue learning. Specifically, the tracking model pre-trained based on the RGB dataset is first used as the base model. Then, a multimodal cue learning module is constructed to learn complementary modal features for effective visual cues.
[0005] Recently popular high-performance RGB-T trackers are mainly divided into two categories. The first type of tracker uses the attention mechanism of the Transformer architecture to construct a multimodal fusion module, which allows cross-modal interaction between RGB and TIR, and fuses single-modal information into multimodal representation, thereby achieving robust tracking; the second type of tracker achieves efficient tracking performance through a multimodal tracking framework based on visual cue learning. Specifically, the tracking model pre-trained based on the RGB dataset is first used as the base model. Then, a multimodal cue learning module is constructed to learn complementary modal features for effective visual cues.
[0006] The above RGB-T tracking methods (the input information is only visual information) can be classified as vision-based RGB-T tracking methods. Although these methods perform well in tracking, they perform poorly in the face of challenges such as camera motion, motion blur, and scale changes. This is because the position, appearance, and shape of the target in these cases will change dramatically, and it becomes difficult to understand the target changes through visual representation alone. Summary of the invention
[0007] In view of the problem in the prior art that pure vision-based RGB-T trackers do not perform well when facing challenges such as camera motion, motion blur, scale change, target appearance change, occlusion and background interference, the present invention provides an RGBT tracking method based on target adaptive text-guided visual fusion, which introduces text modal information other than visual images to assist visual information in target tracking, and uses the text information to combine with the target's visual information to improve the quality of visual features, thereby improving the accuracy of tracking results, so as to more effectively cope with challenges such as camera motion, motion blur and scale change.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A RGBT tracking method based on target adaptive text-guided visual fusion, comprising the following steps: S10, using the BLIP model to generate a corresponding target text description for each frame image in the multimodal dataset; S20, randomly sampling the video sequence and the text description to obtain a multimodal image and a corresponding target text description as training data, and inputting the multimodal image and the corresponding target text description into the RGBT tracking model for training; S30, extracting visual features and text features of the target respectively through a visual encoder and a text encoder; S40. Design a target text adaptive enhancement module based on sparse attention mechanism and insert it into the back end of the text encoder to improve the expressiveness of target text features. At the same time, design a multimodal sharing and complementary information prompter and embed it into the visual encoder to enhance multimodal visual features. S50, fusing the text features output by the text branch with the visual features output by the visual branch through a multimodal encoder to obtain a multimodal fusion feature; S60, performing a convolution operation on the multimodal fusion features to obtain a probability score map and a target box of the tracked target, and then calculating the classification loss and regression loss with the true label to constrain the training process, and finally obtaining a trained RGBT tracking model; S70, performing online tracking, loading the trained RGBT tracking model to test the tracking effect.
[0009] Preferably, in step S10, the visible light image in the multimodal data set is input into the BLIP model, and the content scope of the generated target text description is limited by the set prompt words used to describe the target appearance characteristics.
[0010] Preferably, the multimodal dataset adopts the LasHer dataset.
[0011] Preferably, in step S20, random sampling is performed from a training set of the multimodal dataset.
[0012] Preferably, in step S30, the multimodal image is input into the visual encoder of the visual branch to extract the multimodal visual features of the target, and the target text description corresponding to the multimodal image is input into the text encoder of the text branch to extract the text features of the target.
[0013] Preferably, the process of adaptively enhancing the target text features by the target text adaptive enhancement module based on the sparse attention mechanism in step S40 is expressed as follows: In the above formula, F TZ and F TX The target text description of the template and search image pair respectively represents the text features extracted by the input text encoder, represents a linear layer with 3C output channels, [ ; ] represents a concatenation operation, split( ) represents a slicing operation, Q, K, and V represent the query, key, and value of the sparse attention mechanism input, respectively, L is the total text feature length, is a hyperparameter, topk(·) means selecting the first k elements from each row vector of the L×L matrix and returning the position index of the selected elements, TE and TI represent the text feature matrix and position index matrix of L×k size, respectively, the function scatter(·) inserts the elements of the old matrix into the new matrix according to the position index, ZM is a zero matrix of L×L size, Represents the text features output by the target text adaptive enhancement module.
[0014] Preferably, in step S40, the multimodal sharing and complementary information prompter is embedded in each layer of Transformer block in the visual encoder of the RGB modality and the TIR modality.
[0015] Preferably, the multimodal sharing and complementary information prompter is composed of three branches, each of which mainly includes a spatial attention mechanism and a channel attention mechanism. The process of processing features by the spatial attention SA and the channel attention CA is expressed as follows: In the above formula, F represents the input feature of the multimodal shared and complementary information prompter, MaxPool and AvgPool represent the maximum pooling operation and the average pooling operation respectively, MLP represents the multi-layer perceptron, σ represents the activation function, and f 7×7 Represents a two-dimensional convolution with a convolution kernel size of 7×7.
[0016] Preferably, the process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the RGB modality visual encoder branch is expressed as follows: In the above formula, F 1 、F 2 and F 3 They represent RGB visual features, TIR visual features, and the mixed prompt features of RGB modality and TIR modality output by the previous prompt module, respectively. S and W C They represent spatial attention weight and channel attention weight respectively, LN represents layer normalization, S P1 and S P2 Indicates shared visual cue information, M fovea Represents a mask similar to channel-level spatial attention, used to enhance the shared visual information S P1 , C P represents complementary visual cue information, P represents the shared and complementary information cue information output by the multimodal shared and complementary information prompter, represents element-wise multiplication, Represents matrix multiplication.
[0017] Preferably, the process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the TIR modality visual encoder branch is the same as the process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the RGB modality visual encoder branch, wherein only F 1 and F 2 The features can be exchanged.
[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention proposes a two-stage tracking framework based on text-assisted vision, which utilizes the BLIP model to generate text descriptions of targets in images. The quality of visual features is enhanced by using semantic information such as target type, appearance color, motion state, etc. contained in the text description, thereby improving tracking accuracy and robustness. This effectively solves the problem that pure visual trackers in the prior art perform poorly when faced with challenges such as camera motion, motion blur, size change, target appearance change, occlusion, and background interference.
[0019] (2) The present invention alleviates the negative impact of low-quality text on visual features by designing a target text adaptive enhancement module, and further enhances the expressiveness of visual features by utilizing multimodal sharing and complementary information prompters. These two modules work together to improve the performance of the tracker in complex environments, enabling it to more effectively cope with challenges such as camera motion, motion blur, and scale changes, thereby improving the robustness and accuracy of tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention.
[0021] Figure 2 Schematic diagram of the network structure of an embodiment of the present invention.
[0022] Figure 3 Schematic diagram of the network structure of the target text adaptive enhancement module in an embodiment of the present invention.
[0023] Figure 4 Schematic diagram of the network structure of the multimodal sharing and complementary information prompter in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The present invention is further described below in conjunction with the accompanying drawings and examples. The embodiments of the present invention include but are not limited to the following examples.
[0025] like Figures 1 to 4 As shown, the RGBT tracking method based on target adaptive text guided visual fusion includes the following steps: S10. Use the BLIP model to generate a corresponding target text description for each frame of the multimodal dataset LasHer. Specifically, the visible light image in the multimodal dataset is input into the BLIP model, and the content scope of the generated target text description is limited by the set prompt words used to describe the target appearance characteristics. For example, the prompt word can be "Please use a concise sentence to describe the appearance of the target, focusing on its color, shape and unique features."
[0026] S20. Randomly sample video sequences and text descriptions from the training set of the multimodal dataset LasHer, obtain multimodal images and corresponding target text descriptions as training data, and input them into the RGBT tracking model for training.
[0027] S30, respectively extracting the visual features and text features of the target through the visual encoder and the text encoder. Specifically, the multimodal image is input into the visual encoder (such as VIT) of the visual branch to extract the multimodal visual features of the target, and the target text description corresponding to the multimodal image is input into the text encoder (such as RoBERTa) of the text branch to extract the text features of the target. Among them, the visual encoder is initialized with the weights of the pre-trained RGB tracking network OSTrack256, and contains 12 layers of transformer block layers.
[0028] S40, design a target text adaptive enhancement module based on a sparse attention mechanism, insert it into the back end of the text encoder, and in the text branch, input the target text features extracted by the text encoder into the target text adaptive enhancement module to adaptively enhance the text features related to the target and improve the expressiveness of the target text features. Figure 3 , the process of enhancing the target text features can be expressed by the following formula: In the above formula, F TZ and F TX The target text description of the template and search image pair respectively represents the text features extracted by the input text encoder, represents a linear layer with 3C output channels, [ ; ] represents a concatenation operation, split( ) represents a slicing operation, Q, K, and V represent the query, key, and value of the sparse attention mechanism input, respectively, L is the total text feature length, is a hyperparameter, topk(·) means selecting the first k elements from each row vector of the L×L matrix and returning the position index of the selected elements, TE and TI represent the text feature matrix and position index matrix of L×k size, respectively, the function scatter(·) inserts the elements of the old matrix into the new matrix according to the position index, ZM is a zero matrix of L×L size, Represents the text features output by the target text adaptive enhancement module.
[0029] At the same time, a multimodal shared and complementary information prompter is designed and embedded in the visual encoder to enhance the multimodal visual features. Specifically, in the visual branch, the multimodal shared and complementary information prompter is embedded in each layer of the Transformer block in the visual encoder of the RGB modality and the TIR modality to enhance the expression of visual features. This design aims to promote each modality to jointly learn similar and complementary features from the features of the complementary modality, thereby enhancing the representation ability of visual features. Among them, the network structure diagram of the multimodal shared and complementary information prompter can be found in Figure 4 , the network consists of three branches, each of which mainly includes spatial attention mechanism and channel attention mechanism. The process of spatial attention SA and channel attention CA processing features can be expressed by the following formula: In the above formula, F represents the input feature of the multimodal shared and complementary information prompter, MaxPool and AvgPool represent the maximum pooling operation and the average pooling operation respectively, MLP represents the multi-layer perceptron, σ represents the activation function, and f 7×7 Represents a two-dimensional convolution with a convolution kernel size of 7×7.
[0030] The process of obtaining prompt features by the multimodal sharing and complementary information prompter of the RGB modality visual encoder branch can be expressed by the following formula: In the above formula, F 1 、F 2 and F 3They represent RGB visual features, TIR visual features, and the mixed prompt features of RGB modality and TIR modality output by the previous prompt module, respectively. S and W C They represent spatial attention weight and channel attention weight respectively, LN represents layer normalization, S P1 and S P2 Indicates shared visual cue information, M fovea Represents a mask similar to channel-level spatial attention, used to enhance the shared visual information S P1 , C P represents complementary visual cue information, P represents the shared and complementary information cue information output by the multimodal shared and complementary information prompter, represents element-wise multiplication, Represents matrix multiplication.
[0031] The process of obtaining the prompt feature by the multimodal shared and complementary information prompter of the TIR modality visual encoder branch is the same as the above process, in which only F 1 and F 2 The features can be exchanged.
[0032] S50. Through a multimodal encoder, the text features output by the text branch are fused with the visual features output by the visual branch to obtain a multimodal fusion feature.
[0033] S60, performing a convolution operation on the multimodal fusion features, inputting the multimodal fusion features into the classification regression head composed of the convolution operation, obtaining the probability score map and the target box of the tracked target, and then calculating the classification loss and regression loss with the true label to constrain the training process, and finally obtaining the trained RGBT tracking model.
[0034] S70, performing online tracking, loading the trained RGBT tracking model to test the tracking effect.
[0035] In this embodiment, the designed RGBT cue tracking network is trained on the training set for 12 rounds with a learning rate of 0.00075. Starting from the 10th round, the learning rate decays to 0.000075.
[0036] Table 1 shows the comparison results of the proposed method with some existing RGBT tracking methods on the LasHer and RGBT234 test sets. To ensure the fairness of the comparison, all methods are initialized with the same upstream RGB tracking network pre-trained parameters. The performance evaluation on the LasHeR dataset uses the precision rate (PR), normalized precision rate (NPR), and success rate (SR) indicators, while the evaluation on the RGBT234 dataset uses the PR and SR indicators. By convention, the threshold is set to 20 pixels to calculate a representative PR score.
[0037] Table 1 Comparison of the proposed method with similar methods on the LasHer and RGBT234 test sets The above embodiments are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any changes made by adopting the design principles of the present invention and performing non-creative work on this basis should fall within the protection scope of the present invention.
Claims
1. A RGBT tracking method based on target adaptive text-guided visual fusion, characterized in that: The following steps are involved: S10, using the BLIP model to generate a corresponding target text description for each frame image in the multimodal dataset; S20, randomly sampling the video sequence and the text description to obtain a multimodal image and a corresponding target text description as training data, and inputting the multimodal image and the corresponding target text description into the RGBT tracking model for training; S30, extracting visual features and text features of the target respectively through a visual encoder and a text encoder; S40. Design a target text adaptive enhancement module based on sparse attention mechanism and insert it into the back end of the text encoder to improve the expressiveness of target text features. At the same time, design a multimodal sharing and complementary information prompter and embed it into the visual encoder to enhance multimodal visual features. S50, fusing the text features output by the text branch with the visual features output by the visual branch through a multimodal encoder to obtain a multimodal fusion feature; S60, performing a convolution operation on the multimodal fusion features to obtain a probability score map and a target box of the tracked target, and then calculating the classification loss and regression loss with the true label to constrain the training process, and finally obtaining a trained RGBT tracking model; S70, performing online tracking, loading the trained RGBT tracking model to test the tracking effect.
2. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 1 is characterized in that: In the step S10, the visible light image in the multimodal data set is input into the BLIP model, and the content scope of the generated target text description is limited by the set prompt words used to describe the target appearance characteristics.
3. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 1, characterized in that: The multimodal dataset adopts the LasHer dataset.
4. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 1, characterized in that: In the step S20, random sampling is performed from a training set of the multimodal dataset.
5. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 1, characterized in that: In step S30, the multimodal image is input into the visual encoder of the visual branch to extract the multimodal visual features of the target, and the target text description corresponding to the multimodal image is input into the text encoder of the text branch to extract the text features of the target.
6. The RGBT tracking method based on target adaptive text-guided visual fusion according to any one of claims 1 to 5, characterized in that: The process of the target text adaptive enhancement module based on the sparse attention mechanism in step S40 for adaptively enhancing the target text features is expressed as follows: In the above formula, F TZ and F TX The target text description of the template and search image pair respectively represents the text features extracted by the input text encoder, represents a linear layer with 3C output channels, [ ; ] represents a concatenation operation, split( ) represents a slicing operation, Q, K, and V represent the query, key, and value of the sparse attention mechanism input, respectively, L is the total text feature length, is a hyperparameter, topk(·) means selecting the first k elements from each row vector of the L×L matrix and returning the position index of the selected elements, TE and TI represent the text feature matrix and position index matrix of L×k size, respectively, the function scatter(·) inserts the elements of the old matrix into the new matrix according to the position index, ZM is a zero matrix of L×L size, Represents the text features output by the target text adaptive enhancement module.
7. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 6, characterized in that: In the step S40, the multimodal sharing and complementary information prompter is embedded in each layer of the Transformer block in the visual encoder of the RGB modality and the TIR modality.
8. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 7, characterized in that: The multimodal shared and complementary information prompter consists of three branches, each of which mainly includes a spatial attention mechanism and a channel attention mechanism. The process of processing features by spatial attention SA and channel attention CA is expressed as follows: In the above formula, F represents the input feature of the multimodal shared and complementary information prompter, MaxPool and AvgPool represent the maximum pooling operation and the average pooling operation respectively, MLP represents the multi-layer perceptron, σ represents the activation function, and f 7×7 Represents a two-dimensional convolution with a convolution kernel size of 7×7.
9. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 8, characterized in that: The process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the RGB modality visual encoder branch is expressed as: In the above formula, F1, F2 and F3 represent RGB visual features, TIR visual features and the mixed prompt features of RGB modality and TIR modality output by the previous prompt module, respectively. S and W C They represent spatial attention weight and channel attention weight respectively, LN represents layer normalization, S P1 and S P2 Indicates shared visual cue information, M fovea Represents a mask similar to channel-level spatial attention, used to enhance the shared visual information S P1 , C P represents complementary visual cue information, P represents the shared and complementary information cue information output by the multimodal shared and complementary information prompter, represents element-wise multiplication, Represents matrix multiplication.
10. The RGBT tracking method based on target adaptive text-guided visual fusion according to claim 9, characterized in that: The process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the TIR modality visual encoder branch is the same as the process of obtaining the prompt feature by the multimodal shared and complementary information prompter in the RGB modality visual encoder branch, where only the F1 and F2 features need to be exchanged.
Citation Information
Patent Citations
Single-target tracking method based on multi-mode single-flow memory network
CN116402849A
Feature enhancement multi-modal target tracking method and device based on attention query
CN117333876A
Visual language target tracking method based on deep knowledge distillation
CN118134965A
Content generation method and device, equipment, storage medium and product
CN119763121A
Pre-training of computer vision foundational models
US20230162481A1
Cited By
RGBT target tracking method and system based on interactive hidden state spatio-temporal information
CN120543594A
Structured attribute guided multi-scale context interaction RGB-T sparse fusion tracking method and system
CN121616918A