A target tracking method based on cross-modal feature fusion and refinement of transformer
By introducing a token-adaptive fusion and refinement module, the problems of insufficient cross-modal feature fusion and computational redundancy in RGBT target tracking are solved, achieving efficient cross-modal information interaction and improved target tracking robustness.
Patent Information
- Application Number
- CN202511321894.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing RGBT target tracking methods suffer from insufficient cross-modal feature fusion, computational redundancy, and failure to differentiate between template and search region features, resulting in inadequate robustness and efficiency in complex scenarios.
A token-adaptive fusion and refinement module with asymmetric application is adopted. Through the token-adaptive soft pruning module, the local refinement adapter module, and the cross-modal residual attention module, RGB and infrared search region features are processed respectively to achieve efficient fusion and differentiated treatment.
It improves the robustness and efficiency of the algorithm in complex scenarios, and achieves excellent cross-modal information interaction and accurate background suppression and target focusing capabilities.
Smart Images

Figure CN120852473B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and artificial intelligence, and particularly relates to a target tracking method based on cross-modal feature fusion and refinement of a Transformer. BACKGROUND
[0002] Visual object tracking is a core and challenging task in computer vision, which aims to continuously locate an arbitrary object given in the initial frame in a video sequence. In recent years, with the development of deep learning, especially the great success of the Transformer architecture in the field of computer vision, the tracker based on the Transformer has become mainstream and has shown excellent performance.
[0003] However, when facing complex scenes such as poor environment (such as weak light, strong light, haze) or similar appearance of target and background (such as camouflage), the performance of the tracker relying only on visible light (RGB) images will decrease significantly. In order to solve this problem, infrared and visible light (RGBT) multi-modal tracking emerges as the times require. Infrared thermal imaging does not depend on light and can image according to the thermal radiation of objects, providing effective complementary information for RGB data and significantly improving the robustness of all-weather tracking.
[0004] The core challenge of current RGBT tracking methods is how to effectively and efficiently fuse information from two different modalities. Existing methods usually adopt a simple feature addition or concatenation strategy, which is difficult to capture the deep complex correlation between modalities and the fusion is insufficient. Some methods design complex cross-modal interaction modules, but often introduce a large amount of computational overhead, and in the processing process, the template features and search region features are treated equally, without considering the functional differences between the two in the tracking process (the template provides a stable reference, and the search region contains dynamic changes), causing certain computational redundancy and information interference. Therefore, how to design a more efficient, more sufficient fusion and differentially treat different functional features of the RGBT target tracking method is a direction worth studying and improving in the current field. SUMMARY
[0005] In view of the deficiencies of the prior art, the present application provides a target tracking method based on cross-modal feature fusion and refinement of a Transformer, which solves the problems of insufficient cross-modal feature fusion, computational redundancy and failure to differentially process template and search region features in existing RGBT trackers. By introducing a token self-adaptive fusion and refinement module with asymmetric application, the multi-modal information is effectively fused, and the computational efficiency and tracking robustness are improved.
[0006] To solve the above technical problems, the application provides the following technical scheme: a target tracking method based on cross-modal feature fusion and refinement of a Transformer, comprising the following steps:
[0007] S1, based on the given tracking target initial bounding box, the template image pair and the search image pair in the RGB and infrared video sequence are obtained respectively, the template image pair includes the RGB template and the infrared template, and the search image pair includes the RGB search region and the infrared search region;
[0008] S2, a visual encoder network with a double-flow structure is constructed, the network includes two paths of the same structure and weight sharing Transformer encoders, which are used for extracting and processing the RGB features and the infrared features of the template image pair and the search image pair respectively, to obtain the RGB output features and the infrared output features; a token adaptive fusion and refinement module is asymmetrically inserted after a specific Transformer block of the Transformer encoder;
[0009] S3, the template image pair and the search image pair are input into the visual encoder network with a double-flow structure, the template region features and the search region features are obtained and fused, and the fused feature matrix is input into a prediction head, and finally the position information of the tracking target is output.
[0010] Further, in step S1, the specific process comprises the following steps:
[0011] S11, for the RGB and infrared video sequence, the corresponding regions are cut and normalized as the template image pair according to the given tracking target initial bounding box in the first frame image of the RGB and infrared video sequence respectively;
[0012] S12, in the subsequent t-th frame image , a search box with a larger context range is constructed according to the tracking target center position predicted by the previous frame , and the region corresponding to the search box in the image is cut as the search image pair.
[0013] Further, in step S2, the specific process comprises the following steps:
[0014] S21, the Transformer encoder in the visual encoder network with a double-flow structure includes 12 Transformer blocks, and a token adaptive fusion and refinement module is inserted after the 3rd, 6th and 9th Transformer blocks respectively;
[0015] S22, the token adaptive fusion and refining module is composed of three core sub-modules connected in sequence, including: a token adaptive soft pruning module, a local refining adapter module, and a cross-modal residual attention module;
[0016] S23, the token adaptive soft pruning module, which processes the RGB features and infrared features from a specific Transformer block by probabilizing and constructing a coordinate system, and based on a set of geometric moments calculated independently for each channel, generates an importance weight vector, and then weights the input RGB features and infrared features, outputs RGB weighted features and infrared weighted features;
[0017] S24, the local refining adapter module, which includes a depth separable convolution unit combined with batch normalization and activation function, and a residual connection structure, which enhances the local properties of the features while preserving the original information;
[0018] S25, the cross-modal residual attention module, which uses a multi-head attention mechanism for bidirectional information fusion, and injects the calculated attention features into the input features of the module in the form of residual, to obtain RGB output features and infrared output features , the output relationship can be represented as:
[0019] ;
[0020] ;
[0021] wherein, and respectively represent the RGB weighted features processed by the local refining adapter module and the infrared weighted features processed by the local refining adapter module, represents the standard multi-head attention calculation process, which calculates the attention features.
[0022] Further, in step S23, the specific process includes the following steps:
[0023] S231, for the input RGB features , probabilistic processing is performed on the spatial dimensions of each channel to obtain an information energy density map , and a normalized coordinate system is constructed for the two-dimensional space of the RGB features , that is, a grid with the same spatial dimensions as the RGB features is generated, where each spatial position is assigned a continuous coordinate value ranging from -1 to 1.
[0024] S232. Calculate the RGB features respectively. The average activation value of each channel is used to obtain the zeroth moment, which characterizes the overall signal strength. ;
[0025] S233, Information Energy Density Map Calculate the first moment, i.e., the center of mass, used to accurately determine the location of the center of the information energy space. ;
[0026] S234, Information Energy Density Map Calculate the second-order central distance, or spatial dispersion, used to quantify the distribution pattern of information energy around its center of mass. ;
[0027] S235. Stack the geometric moments of each order calculated in steps S232 to S234 to obtain RGB features. Each channel constructs a high-dimensional geometric descriptor :
[0028] ;
[0029] S236, This high-dimensional geometric descriptor Input a Generate network for parameter weights A nonlinear mapping is performed to analyze its complex spatial morphological information, ultimately yielding an importance weight vector. It can be expressed by the following formula:
[0030] ;
[0031] Among them, the weight generation network The mapping process can be summarized as follows:
[0032] ;
[0033] In the above formula, Represents the compression mapping function. Represents the normalization function. Represents a non-linear activation function, while This represents the regression mapping function;
[0034] S237, the RGB weighted feature finally output by the token adaptive soft pruning module From this importance weight vector Compared with the input RGB features Obtained by element-wise multiplication, the relationship can be expressed as:
[0035] ;
[0036] wherein, The symbol represents a broadcast element-wise multiplication.
[0037] S238, repeat the above steps S231-S237 for the input infrared feature, that is, obtain the infrared weighted feature.
[0038] Further, in step S2, the token adaptive fusion and refinement module adopts an asymmetric processing manner, that is, only the RGB features and infrared features extracted from the search image pair are processed, while the RGB features and infrared features of the template image pair bypass the module without processing, and finally the RGB output features and infrared output features of the search image pair processed by the token adaptive fusion and refinement module are recombined with the RGB output features and infrared output features of the template image pair without processing, and then enter the next Transformer block.
[0039] Further, in step S3, the specific process includes the following steps:
[0040] S31, input the template image pair and the search image pair into the dual-stream structure visual encoder network, and perform feature extraction through multiple Transformer blocks and asymmetrically inserted token adaptive fusion and refinement modules to obtain the final template region features and search region features;
[0041] S32, input the feature matrix obtained after the fusion operation of the template region features and the search region features into the prediction head with a classification branch and a box regression branch, wherein the classification branch is used to judge whether the tracking target exists, and the box regression branch is used to predict the position and size of the tracking target bounding box, and finally output the position information of the tracking target.
[0042] Further, the present application also provides an electronic device, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the target tracking method based on the Transformer cross-modal feature fusion and refinement.
[0043] Further, the present application also provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, wherein the computer program is executed by a processor to implement the steps of the target tracking method based on the Transformer cross-modal feature fusion and refinement.
[0044] By the above technical scheme, the application provides a target tracking method based on cross-modal feature fusion and refinement of a Transformer, and at least has the following beneficial effects:
[0045] (1) The application designs an original token adaptive fusion and refinement module, and through its efficient asymmetric application architecture and structured multi-stage refinement process, excellent cross-modal fusion effect is achieved. The module only refines the search area containing the dynamic tracking target, avoiding redundant calculation of the static template and protecting its reference stability; its internal follows the optimization order of "intra-modal purification, local detail enhancement, and cross-modal complementary mining", ensuring that the final cross-modal information interaction is established on high-quality features, thereby significantly improving the overall robustness and running efficiency of the algorithm in complex scenes;
[0046] (2) The application innovatively designs a new weight evaluation formula in the token adaptive soft pruning module. The module discards the traditional method of relying on feature average activation value and creatively introduces an evaluation mechanism based on computational geometry moments. By calculating the centroid (first moment) and spatial dispersion (second central moment) for each feature channel, the model can obtain a comprehensive geometric fingerprint describing the "spatial form" of the channel feature. This form-based understanding ability enables the model to accurately distinguish between highly concentrated features representing the core components of the target and form-dispersed background noise, even if their average activation values are similar, the former can be given a higher importance weight, thereby achieving unprecedented precise background suppression and target focusing ability. BRIEF DESCRIPTION OF DRAWINGS
[0047] The accompanying drawings used to provide further understanding of the present application and form a part of the present application, illustrate the exemplary embodiments of the present application and the description thereof, and do not constitute improper limitations to the present application. In the drawings:
[0048] Figure 1 The figure is a framework diagram of the target tracking method based on cross-modal feature fusion and refinement of the Transformer of the application;
[0049] Figure 2 The figure is a detailed structure diagram of each sub-module in the token adaptive fusion and refinement module in the application;
[0050] Figure 3 The figure is a success rate and accuracy rate curve diagram of the embodiment method of the application on the LasHeR test data set. DETAILED DESCRIPTION
[0051] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments. In this way, the implementation process of how to apply technical means to solve technical problems and achieve technical effects can be fully understood and implemented.
[0052] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment methods can be completed by programs instructing related hardware, therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program codes.
[0053] Please refer to Figures 1-3 , which shows a specific embodiment of the present embodiment, which solves the problems of insufficient cross-modal feature fusion, computational redundancy and failure to differentiate the features of the template and the search region in the existing RGBT tracker. By introducing a token adaptive fusion and refinement module with asymmetric application, the multi-modal information is effectively fused, and the computational efficiency and the robustness of the tracking are improved.
[0054] The present embodiment proposes a target tracking method based on cross-modal feature fusion and refinement of the Transformer, which comprises the following steps:
[0055] S1, based on the given initial bounding box of the tracking target, the template image pair and the search image pair in the RGB and infrared (RGBT) video sequence are obtained respectively, the template image pair includes the RGB template and the infrared template, and the search image pair includes the RGB search region and the infrared search region;
[0056] As a preferred embodiment of step S1, the specific process comprises the following steps:
[0057] S11, for the RGB and infrared video sequence, the first frame image is cropped and normalized to an image block with a size of 128x128 pixels as the template image pair according to the given initial bounding box of the tracking target , wherein is the center position of the bounding box, is the width and height parameter of the bounding box;
[0058] S12, in the subsequent t-th frame image (wherein ), the previous frame The predicted tracking target center position is used to construct a search frame with a larger context range , and the region corresponding to the search frame is cropped from the image as a search image pair .
[0059] S2, a visual encoder network with a double-flow structure is constructed, and the structure can be referred to Figure 1 The network includes two identical and weight-shared Transformer encoders, which are used to extract and process RGB features and infrared features from the template image pair and the search image pair respectively, to obtain RGB output features and infrared output features. A token adaptive fusion and refinement (TAFR) module is asymmetrically inserted after a specific Transformer block of the Transformer encoder.
[0060] As a preferred embodiment of step S2, the specific process includes the following steps:
[0061] S21, the visual encoder network with a double-flow structure, the Transformer encoder in the network includes 12 Transformer blocks, and the RGB features and infrared features (extracted via the RGB embedding layer and the infrared embedding layer) are input into the Transformer blocks. Token adaptive fusion and refinement (TAFR) modules are inserted after the 3rd, 6th and 9th Transformer blocks respectively.
[0062] S22, the token adaptive fusion and refinement (TAFR) module, and its detailed structure can be referred to Figure 2 It is composed of three core sub-modules connected in order, including:
[0063] A token adaptive soft pruning module (Token-Adaptive Soft Pruner) for soft pruning of features of each modality;
[0064] A local refinement adapter module (Local-Refine Adapter) for local information reconstruction of pruned features;
[0065] A cross-modal residual attention module (Cross-Modal Residual Attention) for bidirectional information interaction and fusion between two modal features;
[0066] S23, the token adaptive soft pruning module, and its specific structure can be referred to Figure 2The module processes and generates an importance weight vector by probabilizing and coordinate system construction of RGB features and infrared features from a specific Transformer block, and based on a set of geometric moments calculated independently for each channel, and then weights the input RGB features and infrared features, and outputs RGB weighted features and infrared weighted features;
[0067] As a preferred embodiment of step S23, the specific process includes the following steps:
[0068] S231, for the input RGB feature , probabilize it in the spatial dimension of each channel by the spatial function Softmax, that is, for each channel , perform exponential operation and normalization on all activation values on the spatial plane, so as to obtain an information energy density map with a dimension of ; Meanwhile, pre-process the two-dimensional space of the RGB feature , and construct a normalized coordinate system, that is, generate a grid with the same spatial dimension as the RGB feature , wherein each spatial position is assigned a continuous coordinate value ranging from -1 to 1;
[0069] S232, calculate the average activation value of each channel of the RGB feature on the grid respectively, to obtain the zero-order moment representing the overall intensity of the signal;
[0070] S233, sum the information energy density map to obtain the first-order moment, i.e. the centroid , for accurately marking the center position of the information energy space, and the calculation method is as follows:
[0071]
[0072] S234, calculate the second-order center distance, i.e. the spatial dispersion , of the information energy density map for quantifying the distribution pattern of the information energy around the centroid, and the calculation method is as follows:
[0073]
[0074] S235, stack the geometric moments calculated in steps S232 to S234 to construct a high-dimensional geometric descriptor containing intensity, position and pattern information for each channel of the RGB feature :
[0075] ;
[0076] S236, the high-dimensional geometric descriptor is input into a weight generation network ; The nonlinear mapping is performed to analyze the complex spatial form information, and finally the importance weight vector can be represented by the following formula:
[0077] ;
[0078] The mapping process of the weight generation network can be summarized as:
[0079] ;
[0080] In the above formula, represents a compression mapping function, represents a normalization function, represents a nonlinear activation function, and represents a regression mapping function;
[0081] S237, the RGB weighted feature finally output by the token adaptive soft pruning module is obtained by element-wise multiplication of the importance weight vector and the input RGB feature , and the relationship can be represented as:
[0082] ;
[0083] wherein, the symbol represents element-wise multiplication by broadcasting;
[0084] S238, the above steps S231-S237 are repeated for the input infrared feature, i.e., the infrared weighted feature is obtained.
[0085] In this embodiment, a brand new weight evaluation formula is designed in the token adaptive soft pruning module. This module discards the traditional method of relying on the average activation value of the feature, and innovatively introduces an evaluation mechanism based on the calculation of geometric moments. By calculating the centroid (first moment) and spatial dispersion (second central moment) for each feature channel, the model can obtain a comprehensive geometric fingerprint that describes the "spatial form" of the channel feature. This form-based understanding ability enables the model to accurately distinguish between highly concentrated features representing the core components of the target and form-dispersed background noise. Even if the average activation values of the two are similar, the former can still be given a higher importance weight, thereby achieving unprecedented precise background suppression and target focusing ability.
[0086] S24, the local refining adapter module, comprising a depth separable convolution unit composed of a 3x3 depth convolution with a hole and a 1x1 point-wise convolution, combined with batch normalization and activation function, to obtain refined features; finally, the refined features are added to the features input into the module through a residual connection, to enhance the local property of the features while preserving the original information;
[0087] S25, the cross-modal residual attention module, which utilizes multi-head attention mechanism for bidirectional information fusion, and injects the calculated attention features into the input features of the module in a residual form through two independent learnable parameters and to obtain RGB output features and infrared output features , the above output relationship can be expressed as:
[0088] ;
[0089] ;
[0090] wherein, and respectively represent the RGB weighted features processed by the local refining adapter module and the infrared weighted features processed by the local refining adapter module, represents the standard multi-head attention calculation process (i.e. Cross Attention in Figure 2 ), which receives queries , keys , and values from the RGB branch (rgb) and from the infrared branch (tir), and calculates attention features.
[0091] More specifically, the token adaptive fusion and refinement (TAFR) module adopts an asymmetric processing method, i.e. only the RGB features and infrared features extracted from the search image pair are processed (i.e. only the search region tokens of the RGB and infrared embedding layers are input), while the RGB features and infrared features of the template image pair bypass the module without processing, and finally the RGB output features and infrared output features of the search image pair processed by the token adaptive fusion and refinement (TAFR) module are recombined with the RGB output features and infrared output features of the template image pair without processing, and are jointly input into the next Transformer block.
[0092] In this embodiment, the designed token adaptive fusion and refinement (TAFR) module realizes excellent cross-modal fusion effect through its efficient asymmetric application architecture and structured multi-stage refinement process. The module only refines the search area containing the dynamic tracking target, avoiding redundant calculation of the static template and protecting its reference stability; its internal follows the optimization order of "intra-modal purification, local detail enhancement, and cross-modal complementary mining", ensuring that the final cross-modal information interaction is based on high-quality features, thereby significantly improving the overall robustness and running efficiency of the algorithm in complex scenes.
[0093] S3, input the template image pair and the search image pair into the visual encoder network with a double-stream structure respectively, obtain the template region feature and the search region feature, and fuse them, input the fused feature matrix into a prediction head, and finally output the position information of the tracking target;
[0094] As a preferred embodiment of step S3, the specific process includes the following steps:
[0095] S31, input the template image pair and the search image pair into the visual encoder network with a double-stream structure, perform feature extraction through multiple layers of Transformer blocks and asymmetrically inserted token adaptive fusion and refinement (TAFR) modules, and obtain the final template region feature and search region feature.
[0096] S32, concatenate or add the template region feature and the search region feature to obtain a feature matrix, and input the feature matrix into a prediction head with a classification branch and a bounding box regression branch, wherein the classification branch is used to judge whether the tracking target exists, and the bounding box regression branch is used to predict the position and size of the tracking target bounding box, and finally output the position information of the tracking target.
[0097] Repeat all the above steps until the last frame of the RGB and infrared video sequence is tracked.
[0098] In this embodiment, the LasHeR dataset is selected as the training and test dataset, and the batchsize size of the model training is set to 16, and in each epoch, 60000 pictures are randomly selected from the dataset. Secondly, the AdmW optimizer is used to update the weights of the visual encoder network with a double-stream structure. Among them, the weight decay coefficient is set to 0.0001, and the total training of the network is performed for 15 epochs. In the first 10 epochs, the fixed setting of the learning rate of the backbone part is , the learning rate of the proposed TAFR module is set to The learning rate is reduced to 10 times of the original after 10 epochs. This learning rate setting strategy can keep the pre-trained network stable while the new module converges quickly in the first 10 epochs, and then fine-tune in the last 5 epochs, which balances the efficient training and the best generalization in the short period of 15 epochs.
[0099] Figure 3 The success rate and accuracy rate curves of the method of the embodiment on the LasHeR test dataset are shown in the figure. In the success rate figure on the upper side, the horizontal coordinate represents the position error threshold (here, the overlap rate threshold) between the predicted bounding box and the real bounding box of the tracked target, and the vertical coordinate is the success rate. The success rate of the method of the embodiment (ours, the red curve in the figure) is 0.555, which is better than 0.409 of the baseline method (baseline, the green curve in the figure). Figure Three In the accuracy rate figure on the lower side, the horizontal coordinate represents the position error threshold (here, the pixel distance threshold) between the center points of the predicted bounding box and the real bounding box of the tracked target, and the vertical coordinate is the accuracy rate. The accuracy rate of the method of the embodiment (ours, the red curve in the figure) is 0.694, which is also better than 0.493 of the baseline method (baseline, the green curve in the figure). Generally, the higher the values of the success rate (the area under the curve) and the accuracy rate, the better the performance of the tracker.
[0100] The embodiment of the present application also provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the computer program is configured to implement the target tracking method based on cross-modal feature fusion and refinement of the Transformer.
[0101] The embodiment of the present application also provides a storage medium, which is a computer readable storage medium, and the storage medium stores a computer program, and the computer program is executed by a processor to implement the target tracking method based on cross-modal feature fusion and refinement of the Transformer.
[0102] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0103] The logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions for implementing logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions.
[0104] The above embodiments introduce the present application in detail, the principles and embodiments of the present application are described by applying specific examples, the above example is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific embodiment and application range will be changed, according to the above, the content of the specification should not be understood as the limitation of the present application.
Claims
1. A target tracking method based on Transformer for cross-modal feature fusion and refinement, characterized in that, Includes the following steps: S1. Based on the given initial bounding box of the tracking target, obtain template image pairs and search image pairs in the RGB and infrared video sequences respectively. The template image pairs include RGB templates and infrared templates, and the search image pairs include RGB search regions and infrared search regions. S2. Construct a dual-stream visual encoder network, which includes two identical Transformer encoders with shared weights, used to extract and process RGB and infrared features from template image pairs and search image pairs respectively, to obtain RGB output features and infrared output features. The Transformer encoder is followed by a token adaptive fusion and refinement module that is asymmetrically inserted after a specific Transformer block; S21. The dual-stream visual encoder network, wherein the Transformer encoder includes 12 Transformer blocks, and token adaptive fusion and refinement modules are inserted after the 3rd, 6th and 9th Transformer blocks respectively. S22. The token adaptive fusion and refinement module is composed of three core sub-modules connected in sequence, including: token adaptive soft pruning module, local refinement adapter module, and cross-modal residual attention module; S23. The token adaptive soft pruning module, which probabilizes and constructs a coordinate system for RGB and infrared features from a specific Transformer block, and processes them based on a set of geometric moments calculated independently for each channel to generate an importance weight vector, and then weights the input RGB and infrared features to output RGB weighted features and infrared weighted features. S231. For the input RGB features Information energy density maps are obtained by performing probabilistic processing on the spatial dimension of each channel. At the same time, it is the RGB feature A normalized coordinate system is constructed in the two-dimensional space, that is, a coordinate system corresponding to the RGB feature is generated. A grid with the same spatial dimensions, where each spatial location They are all assigned a continuous coordinate value in the range of -1 to 1; S232. Calculate the RGB features respectively. The average activation value of each channel is used to obtain the zeroth moment, which characterizes the overall signal strength. ; S233, Information Energy Density Map Calculate the first moment, i.e., the center of mass, used to accurately determine the location of the center of the information energy space. ; S234, Information Energy Density Map Calculate the second-order central distance, or spatial dispersion, used to quantify the distribution pattern of information energy around its center of mass. ; S235. Stack the geometric moments of each order calculated in steps S232 to S234 to obtain RGB features. Each channel constructs a high-dimensional geometric descriptor : ; S236, This high-dimensional geometric descriptor Input a Weights for parameters in a network A nonlinear mapping is performed to analyze its complex spatial morphological information, ultimately yielding an importance weight vector. It can be expressed by the following formula: ; Among them, the weight generation network The mapping process can be summarized as follows: ; In the above formula, Represents the compression mapping function. Represents the normalization function. Represents a non-linear activation function, while This represents the regression mapping function; S237, the RGB weighted feature finally output by the token adaptive soft pruning module From this importance weight vector Compared with the input RGB features Obtained by element-wise multiplication, the relationship can be expressed as: ; in, The symbol represents element-wise multiplication in broadcast form; S238. Repeat steps S231-S237 above for the input infrared features to obtain the infrared weighted features. S24. The local refinement adapter module includes a depthwise separable convolutional unit that combines batch normalization and activation functions, and a residual connection structure that preserves the original information while enhancing feature locality. S25. The cross-modal residual attention module utilizes a multi-head attention mechanism for bidirectional information fusion and employs two independent learnable parameters. and The calculated attention features are injected as residuals into the input features of the back-to-base module to obtain the RGB output features. and infrared output characteristics The above output relationship can be expressed as: ; ; in, and These represent the RGB weighted features and the infrared weighted features after processing by the local refinement adapter module, respectively. This represents the standard multi-head attention computation process, which calculates attention features. S3. Input the template image pair and the search image pair into the dual-stream visual encoder network respectively, obtain the template region features and the search region features, and fuse them. Then, send the fused feature matrix into the prediction head and finally output the position information of the tracking target.
2. The target tracking method based on Transformer for cross-modal feature fusion and refinement as described in claim 1, characterized in that: Step S1 specifically includes the following steps: S11. For the RGB and infrared video sequences, respectively, in their first frame image... In this process, based on the given initial bounding box of the tracking target, the corresponding region is cropped and normalized as a template image pair; S12, in the subsequent t-th frame image In the middle, based on the previous frame The predicted center location of the tracking target is used to construct a search box with a wider context, and from the image... The region corresponding to the search box is cropped out as the search image pair.
3. The target tracking method based on Transformer cross-modal feature fusion and refinement according to claim 1, characterized in that: The token adaptive fusion and refinement module adopts an asymmetric processing method, that is, it only processes the RGB and infrared features extracted from the search image pair, while the RGB and infrared features of the template image pair are bypassed by the module and not processed. Finally, the RGB and infrared output features of the search image pair processed by the token adaptive fusion and refinement module are recombined with the RGB and infrared output features of the unprocessed template image pair and enter the next Transformer block together.
4. The target tracking method based on Transformer for cross-modal feature fusion and refinement according to claim 3, characterized in that: Step S3 specifically includes the following steps: S31. The template image pair and the search image pair are fed into a two-stream visual encoder network. After feature extraction through a multi-layer Transformer block and an asymmetric insertion token adaptive fusion and refinement module, the final template region features and search region features are obtained. S32. The feature matrix obtained by fusing the template region features and the search region features is input into the prediction head with a classification branch and a bounding box regression branch. The classification branch is used to determine whether the tracking target exists, and the bounding box regression branch is used to predict and adjust the position and size of the tracking target's bounding box. Finally, the position information of the tracking target is output.
5. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the target tracking method based on Transformer cross-modal feature fusion and refinement as described in any one of claims 1 to 4.
6. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the target tracking method based on Transformer cross-modal feature fusion and refinement as described in any one of claims 1 to 4.
Citation Information
Patent Citations
RGBT target tracking method based on target perception enhancement fusion structure
CN117474957A
RGBT tracking method based on CNN-Transform mixing
CN120147357A