Modal sharing information layered unwrapping fusion network for RGB-T target tracking
Through the modal shared information layered de-entanglement fusion network, the ResNet-50 network and cross-modal attention mechanism are used to solve the problem of insufficient utilization of modal residual information in RGB-T fusion tracking, and efficient tracking in complex environments is achieved, and robustness and accuracy are improved.
Patent Information
- Application Number
- CN202510819494.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In the existing RGB-T fusion tracking method, the modal residual information is insufficiently utilized, the fusion method is rough, and the tracking stability is poor, especially in occlusion, low illumination and complex backgrounds.
The modal shared information layered detangling fusion network is adopted, including the ResNet-50 network with parameter sharing, a cross-modal attention module, a hierarchical detangling mining module and an adaptive fusion module. The deep complementary information between the mining modes is calculated through the bidirectional attention mechanism and multi-level residuals, and multi-modal feature fusion is realized through dynamic weight allocation.
It significantly improves the utilization rate of modal complementary information, enhances the robustness and stability of the system in complex environments, solves the performance degradation caused by cumulative errors in long sequences, and improves tracking accuracy.
Smart Images

Figure CN120339780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-modal visual tracking, and specifically to a hierarchical disentanglement and fusion network for modality-shared information in RGB-T object tracking. Background Art
[0002] With the rapid development of multi-modal perception technology, fusing visible light (RGB) images with thermal infrared images (Thermal) for object tracking is becoming a research hotspot in the field of visual perception. Since RGB images are sensitive to the texture and color of objects, while thermal infrared images have stronger object identification capabilities in low-light or occluded scenarios, and the two are highly complementary in the perception mechanism, the RGB-T fusion tracking technology is widely used in high-demand scenarios such as unmanned aerial vehicle monitoring, night security patrols, and traffic monitoring under complex weather conditions. Current research mainly focuses on how to effectively integrate the two modalities of information to improve the robustness and discrimination ability of object representation, especially to obtain accurate and stable tracking results under challenging conditions such as occlusion, low illumination, and complex backgrounds.
[0003] However, there are still many limitations in the modality information fusion methods of existing RGB-T fusion tracking methods. Most methods use a unified attention mechanism to perform a single interaction on modality features. Although they perform well in the initial frame, in the face of scenes with large inter-frame differences or inconsistent modality responses, it is extremely easy to ignore key but inconspicuous complementary residual information, resulting in incomplete modeling of the object appearance by the model. In addition, many methods directly perform feature splicing or simple weighted fusion in the fusion stage. Although it is easy to implement, it is easy to introduce redundant or conflicting information, making it difficult for the model to effectively capture fine-grained modality differences, thereby affecting the tracking accuracy. On the other hand, some methods also try to separately model modality branches in the late fusion decision layer, but such methods often rely on the modeling quality of the modality itself. Once the quality of a certain modality image deteriorates, such as the RGB image being overexposed due to strong light or the thermal map being blurred, the tracking effect after fusion will degrade significantly. In addition, existing methods generally rely on historical frame modeling. Especially in long sequences, this cumulative mechanism will generate irreversible drift errors when the object is briefly occluded or severely deformed.
[0004] In summary, although the RGB-T fusion tracking technology has made great progress in both the academic and engineering fields, the existing technology has not been able to fully explore the hidden weak matching information between modalities, and the fusion strategy still tends to be static or rough, lacking the ability to finely model multi-scale and hierarchical complementary relationships. Therefore, the present invention proposes a hierarchical disentanglement and fusion network for modality-shared information in RGB-T object tracking. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a hierarchical disentangling and fusion network for RGB-T target tracking, which solves the problems of insufficient utilization of modal residual information, rough fusion method, and poor tracking stability in the existing RGB-T fusion tracking methods.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A hierarchical disentangling and fusion network for RGB-T target tracking, including: A dual-stream feature extraction module, which uses a parameter-sharing ResNet-50 network to extract the features of RGB images and thermal infrared images respectively; A cross-modal attention module, which is connected to the dual-stream feature extraction module to enhance the feature interaction between RGB and thermal modalities through a bidirectional attention mechanism; A hierarchical disentangling and mining module, which is connected to the cross-modal attention module to mine the deep complementary information between modalities through multi-level residual calculation; An adaptive fusion module, which is connected to the hierarchical disentangling and mining module to achieve multi-modal feature fusion based on dynamic weight allocation; A tracking prediction module, which is connected to the adaptive fusion module to output the target position and scale information.
[0007] Preferably, the cross-modal attention module includes: A thermal modality enhancement path, using the thermal feature as the query vector and the RGB feature as the key-value pair, and calculating the attention weight through formula one; An RGB modality enhancement path, using the RGB feature as the query vector and the thermal feature as the key-value pair, and calculating the attention weight through formula two.
[0008] Preferably, the specific formula one in the thermal modality enhancement path is: ; Where is the cross-modal attention calculation function, represents the Softmax function, is the query vector generated by the linear projection of the thermal modality feature, is the key vector generated by the linear projection of the modality feature, is the value vector generated by the linear projection of the modality feature, ; Where is the query vector generated by the linear projection of the RGB modality feature, The key vector generated by linear projection of thermal modality features The value vector generated by linear projection of thermal modality features is the cross-modal attention calculation function
[0009] Preferably, the hierarchical disentanglement mining module performs the following operations Iteratively calculate the residual features, and layer by layer strip the exploited complementary information through Formula 3 and Formula 4 Accumulate multi-level attention features, and aggregate the deep complementary information through Formula 5 and Formula 6
[0010] Preferably, Formula 3 and Formula 4 in the iterative calculation of residual features are specifically ; ; where and are the disentangled features of RGB and thermal infrared at the th layer respectively is the query vector generated by the RGB features of the th layer is the key vector generated by the RGB features of the th layer is the value vector generated by the RGB features of the th layer is the key vector generated by the thermal features of the th layer is the value vector generated by the thermal features of the th layer is the disentanglement level index is the RGB residual disentanglement output feature of the th layer is the thermal modality residual disentanglement output feature of the th layer
[0011] Preferably, Formula 5 and Formula 6 in the accumulation of multi-level attention features are specifically ; ; where is the total number of disentanglement levels is the query vector generated by the thermal features of the th layer is the cross-modal attention calculation function is the cumulative complementary feature from RGB to thermal modality is the cumulative complementary feature from thermal to RGB modality
[0012] Preferably, the adaptive fusion module includes: A feature evaluation unit that calculates the modal weight through Equation 7; A probability weighting unit that realizes feature fusion through Equation 8.
[0013] Preferably, the specific form of Equation 7 is: ; Where is the enhanced RGB / thermal modal feature, is the global average pooling, is the global max pooling, is the channel concatenation operation, is the fused feature description vector, is the multi-layer perceptron, is the hyperbolic tangent activation function, is the rectified linear unit, is the modal weight coefficient; The specific form of Equation 8 is: ; Where is the fusion weight coefficient of the RGB modality, is the fusion weight coefficient of the thermal modality, is the enhanced RGB feature, is the enhanced thermal modal feature, is the final fused feature.
[0014] Preferably, the tracking and prediction module includes: A classification branch that predicts the target presence probability map; A regression branch that outputs the target bounding box parameters; A decoding unit that converts the regression parameters into (x, y, w, h) coordinates.
[0015] The present invention provides a modality-shared information hierarchical disentanglement fusion network for RGB-T target tracking. It has the following beneficial effects: 1. By introducing a multi-scale decoupling structure, the present invention explicitly separates and dynamically fuses the features of different modalities, achieving the technical effect of significantly improving the utilization rate of modality complementary information. Compared with the existing methods of directly concatenating or simply weighted fusing multi-modalities, this solution effectively solves the problems of information redundancy and weak matching signals being masked.
[0016] 2. The present invention adopts a ToMP-based tracking framework and integrates a lightweight attention module and a modality alignment mechanism to achieve fine-grained correction of potential semantic differences between thermal infrared and visible light images. In complex environments, its tracking performance is stable. Compared with existing unaligned modeling technologies, the robustness in low-light and occlusion scenarios is significantly improved.
[0017] 3. The present invention adopts an end-to-end training strategy and is fully verified on large-scale challenge datasets such as LasHeR and VTUAV. It surpasses existing mainstream algorithms in most key metrics. Compared with traditional schemes that rely on historical frame modeling, this method can obtain better tracking performance by only using the initial frame and the current frame, solving the problem of performance degradation caused by cumulative errors in long sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is an architecture diagram of the modality-shared information hierarchical disentanglement and fusion tracking network of the present invention, where CA extracts significant complementary information, HDM mines residual weak complementary information, and AF integrates multi-modal features; Figure 2 It is a flowchart of the module framework of the system of the present invention; Figure 3 It is a framework diagram of the cross-modal attention module of the present invention; Figure 4 It is a framework diagram of the adaptive fusion module of the present invention; Figure 5 It is a framework diagram of the tracking prediction module of the present invention; Figure 6 It is a schematic diagram of the performance evaluation of the present invention on GTOT, RGBT234, and LasHeR; Figure 7 It is a schematic diagram of the performance evaluation of the present invention on VTUAV; Figure 8 It is an enlarged comparison of the bike-5 scenario in the visualization comparison diagram of the tracking results of the present invention Figure 1 ; Figure 9 It is an enlarged comparison of the bike-5 scenario in the visualization comparison diagram of the tracking results of the present invention Figure 2 ; Figure 10 It is an enlarged comparison of the bike-5 scenario in the visualization comparison diagram of the tracking results of the present invention Figure 3 ; Figure 11 It is an enlarged comparison of the car-72 scenario in the visualization comparison diagram of the tracking results of the present invention Figure 1 ; Figure 12 It is an enlarged comparison of the car-72 scenario in the visualization comparison diagram of the tracking results of the present invention Figure 2 ; Figure 13 The enlarged comparison of the car-72 scenario in the visualization comparison graph of the tracking results of the present invention Figure 3 ; Figure 14 The enlarged comparison of the excavator-001 scenario in the visualization comparison graph of the tracking results of the present invention Figure 1 ; Figure 15 The enlarged comparison of the excavator-001 scenario in the visualization comparison graph of the tracking results of the present invention Figure 2 ; Figure 16 The enlarged comparison of the excavator-001 scenario in the visualization comparison graph of the tracking results of the present invention Figure 3 ; Figure 17 The enlarged comparison of the pedestrian-164 scenario in the visualization comparison graph of the tracking results of the present invention Figure 1 ; Figure 18 The enlarged comparison of the pedestrian-164 scenario in the visualization comparison graph of the tracking results of the present invention Figure 2 ; Figure 19 The enlarged comparison of the pedestrian-164 scenario in the visualization comparison graph of the tracking results of the present invention Figure 3 。 Detailed implementation manners
[0019] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Please refer to the attached Figure 1 - attached Figure 2 , the embodiment of the present invention provides a modality-shared information hierarchical disentanglement and fusion network for RGB-T target tracking, including: A dual-stream feature extraction module, which uses a parameter-sharing ResNet-50 network to extract features of RGB images and thermal infrared images respectively; The dual-stream feature extraction module includes a ResNet-50 network structure with parameter sharing. Its first stream processes the RGB image input, and the second stream processes the thermal infrared image input. The two streams of the network share the weight parameters of the first four convolutional blocks. Exemplarily, the input image size is adjusted to 320×320 pixels, and the output feature map size is 20×20×1024 after being processed by the convolutional blocks.
[0021] During specific implementation, the RGB image input unit first performs normalization processing: ; Among them, is the mean value of the RGB channels of the ImageNet dataset; is the standard deviation; is the original RGB input image, with dimensions H×W×3; is the normalized RGB image, with a numerical range of approximately [-2.12, 2.64]. The thermal infrared image input unit is converted into a three-channel pseudo-color through gray-scale replication: ; Among them, is the original thermal infrared input image, with a single-channel dimension of H×W×1; is the replication operation along the channel axis, expanding the single channel into three channels; is the pseudo-three-channel thermal image, with dimensions H×W×3.
[0022] The two-stream network extracts features through the same convolution operation: ; ; ; ; Among them, is the first convolution block, containing 1 7×7 convolution layer (stride 2), with 64 output channels and a feature map size of 160×160; is the second convolution block, containing 3 residual units, with 256 output channels and a feature map size of 80×80; is the third convolution block, containing 4 residual units, with 512 output channels and a feature map size of 40×40; is the fourth convolution block, containing 6 residual units, with 1024 output channels and a feature map size of 20×20; is the RGB feature map output at the stage; is the final RGB feature map, with dimensions 20×20×1024.
[0023] The calculation process of the thermal flow feature is the same as that of the RGB stream, sharing the convolution kernel parameters of Conv1 to Conv4. The convolution block contains a residual connection structure, which can alleviate the problem of gradient disappearance.
[0024] Preferably, the size of the convolution kernel of Conv1 is 7×7, the stride is 2, and the number of output channels is 64; Conv2 contains 3 residual units, and the number of output channels is 256; Conv3 contains 4 residual units, and the number of output channels is 512; Conv4 contains 6 residual units, and the number of output channels is 1024. Each residual unit contains a bottleneck structure, and the number of channels is adjusted by 1×1 convolution.
[0025] In terms of technical effects, the parameter sharing mechanism can reduce the number of model parameters by about 50%, and at the same time force the bimodal features to be aligned in the same latent space, providing a basis for subsequent cross-modal interaction. The design of gradually reducing the size of the feature map is beneficial to extracting high-level semantic features while retaining spatial information.
[0026] Please refer to the attached Figure 3 , a cross-modal attention module, which is connected to the two-stream feature extraction module and realizes the enhancement of feature interaction between the RGB and thermal modalities through a bidirectional attention mechanism; The present invention provides a cross-modal attention module. The module is used to establish an information interaction mechanism between the RGB modality and the thermal infrared modality to enhance the bimodal joint representation ability. The module structure includes a feature projection part, an attention calculation part and a feature fusion part. The module is logically connected after the two-stream feature extraction module and physically receives the feature maps from the RGB and thermal modalities as inputs.
[0027] The role of the feature projection part is to project the input RGB modality features and thermal modality features into query, key, and value vectors respectively. Let be the number of spatial positions after the feature map is flattened, and be the original channel dimension.
[0028] In the thermal modality enhancement path, the linear projection is calculated as follows: ; ; ; Among the above variables, represents the query vector generated by the linear projection of the thermal modality features, and are the key and value vectors generated by the RGB modality features, is the linear transformation matrix for mapping the thermal modality features into query vectors, is the linear transformation matrix for mapping the RGB modality features into key vectors, is the linear transformation matrix for mapping the RGB modality features into value vectors, is the original input thermal infrared modality feature, is the original RGB modal feature input. and is the attention dimension set to reduce the computational amount, satisfying .
[0029] The attention calculation unit performs the following operations to achieve cross-modal enhancement of the RGB modality to the thermal modality: ; Among them, represents the Softmax function, and the normalization direction is the dimension of the key vector. is the scaling factor used to suppress gradient explosion and stabilize training. is to the complementary feature of the thermal image. is the key vector generated by linear projection of the modal feature. is the value vector generated by linear projection of the modal feature. is the query vector generated by linear projection of the thermal modal feature.
[0030] The way to fuse and enhance the result is to add it to the original thermal modal feature to form a residual connection: ; The symmetric structure is the RGB modality enhancement path. First, construct the query, key, and value vectors as follows: ; ; ; Among them, is the linear transformation matrix that maps the RGB modal feature to the query vector. is the linear transformation matrix that maps the thermal modal feature to the key vector. is the linear transformation matrix that maps the thermal modal feature to the value vector. is the query vector generated by linear projection of the RGB modal feature. is the key vector generated by linear projection of the thermal modal feature. is the value vector generated by linear projection of the thermal modal feature.
[0031] Execute the attention calculation of the RGB modality enhancement path as follows: ; Among them, is the query vector generated by linear projection of the RGB modal feature. is the key vector generated by linear projection of the thermal modal feature. The value vector generated by the linear projection of the thermal modal features represents the Softmax function is the scaling factor is the cross-modal attention calculation function
[0032] The fused RGB modal enhanced features are as follows ; All the linear mapping weight matrices in the above two paths are not shared, and the specific constraints are as follows ; Each group of weight matrices is independently learned to ensure the expression independence of the two cross-modal paths and avoid feature cross-interference
[0033] The implementation method of the feature fusion part is pointwise addition, logically connecting the projection part and the attention output end. The and output by the module are used as the input of the subsequent network
[0034] The Softmax function normalization process acts on the key vector direction corresponding to each query vector to ensure that the sum of the attention weights is 1, meeting the requirements of the attention mechanism
[0035] In an example, assume that the size of the feature map is 20×20, that is, n = 400. Then each attention calculation involves the calculation of a 400×400 similarity matrix and outputs dimensional cross-modal fusion features
[0036] The initialization strategies of each projection matrix are as follows Use He initialization and Use Xavier initialization to balance the variance in the forward propagation and ensure the numerical stability of the training process
[0037] The execution process of the module is as follows: First, the dual-stream feature extraction module obtains and , and then inputs them into the cross-modal attention module. Projection, attention calculation and residual fusion are respectively completed in the thermal modal enhancement path and the RGB modal enhancement path, and finally the features after dual-modal fusion are output
[0038] In the specific implementation, the feature extraction module can be a shared backbone structure, such as a ResNet branch; the module can be deployed in the middle of the backbone network or at the multi-scale fusion node to achieve cross-modal feature collaborative enhancement
[0039] The hierarchical disentanglement and mining module, which is connected to the cross-modal attention module, mines the deep complementary information between modalities through multi-level residual calculations This embodiment provides a hierarchical disentanglement and mining module, which is used to further extract deep residual complementary information from the features output by the cross-modal attention mechanism, so as to improve the robustness and detailed expression ability of cross-modal information fusion.
[0040] The module includes a residual disentanglement part and a multi-level aggregation part, and its structure is connected to the output end of the cross-modal attention module. The input is and two modal complementary enhanced features, and the output is a depth-enhanced feature and .
[0041] First, a disentangled feature initialization model is established, and the original input feature of the 0th layer is defined as follows: ; When , the initial disentangled feature is set as follows: ; Among them, , respectively represent the disentangled input features of the RGB and thermal modalities of the th layer.
[0042] In the residual disentanglement stage, the strongly complementary information that has been extracted is stripped through layer-by-layer Attention operation, and the fine-grained modal residual information is retained. The specific process is as follows: ; ; Among them, is the RGB residual disentanglement output feature of the th layer, is the thermal modality residual disentanglement output feature of the th layer, and are the disentangled features of RGB and thermal infrared at the th layer respectively, is the query vector generated by the RGB feature of the th layer, is the key vector generated by the RGB feature of the th layer, is the value vector generated by the RGB feature of the th layer, is the key vector generated by the thermal feature of the th layer, is the value vector generated by the thermal feature of the th layer, is the disentanglement level index; The query, key, and value vectors involved in the above calculations are defined as follows: ; ; ; ; Among them, is the query vector generated by the -th layer of thermal feature, is the RGB query weight of the -th layer, is the thermal infrared query weight of the -th layer, is the RGB key weight of the -th layer, is the thermal modality key weight of the -th layer, is the thermal modality value weight of the -th layer, is the linear mapping weight of the -th layer of features, with dimensions as follows: ; ; The attention calculation function is defined as follows: ; Among them, is the Softmax function, used to stabilize the gradient, is the scaled dot-product attention output based on .
[0043] The above residual unwrapping process can be iteratively executed layers to retain deeper complementary details, and then enter the multi-level feature aggregation stage after completion.
[0044] The multi-level aggregation part accumulates the complementary detailed features extracted from all levels, and the aggregation process is as follows: ; ; Among them, and respectively represent the complementary enhanced information aggregated from the RGB modality and thermal modality unwrapping processes, is the total number of unwrapping levels, is the query vector generated by the -th layer of thermal feature, is the cross-modal attention calculation function.
[0045] The finally output deep enhanced features are combined by the original fused features and the above complementary information, and are expressed as follows: ; ; Among them, is the finally enhanced RGB modality feature, is the finally enhanced thermal infrared modality feature.
[0046] After the above operations are completed, the output , is the depth-enhanced modality feature that can finally be used for downstream tasks such as detection, segmentation, or recognition.
[0047] The weight parameters of each layer of the module are not shared to ensure the independence of the representations between levels. The linear mapping part adopts He initialization, and adopts Xavier initialization. The attention calculation uses the scaled dot product mechanism and is normalized by Softmax to ensure numerical stability.
[0048] The module structure includes: Residual unwrapping part: Connecting the output of the cross-modal attention, performing the calculation formula and the calculation formula to strip redundant complementary information Multi-level aggregation part: Connecting the output of the unwrapping part, performing the calculation formula and the calculation formula to complete the cross-level complementary information fusion The overall logical process of the module is: First, input the fused features from the output of the cross-modal attention module and ; Subsequently, through the unwrapping iteration mechanism, perform residual calculations layer by layer to strip shallow information and extract residual complementary information; Finally, cumulatively aggregate the attention features extracted from all levels and output the enhanced features and .
[0049] This module can be integrated into the middle layer of the YOLO series detectors, Transformer encoder structures, or FPN feature fusion networks. Appropriately setting the number of layers (recommended value 2 - 4) can achieve a trade-off between computational efficiency and representational ability.
[0050] The hierarchical unwrapping and mining module realizes the layer-by-layer stripping and depth enhancement of the residual complementary information between modalities, can significantly improve the fine-grained feature expression ability of cross-modal fusion representations, and further enhance the performance of multi-modal perception tasks such as object detection and scene understanding.
[0051] Please refer to the appendix Figure 4 , an adaptive fusion module, which is connected to the hierarchical disentangling and mining module and realizes multimodal feature fusion based on dynamic weight allocation; This embodiment provides an adaptive fusion module, which is applicable to the RGB and thermal infrared modality fusion tasks, and improves the information complementary extraction ability and fusion robustness.
[0052] The system includes: a cross-modal attention module, a hierarchical disentangling and mining module, and an adaptive fusion module. The physical connections of each module are as follows: The cross-modal attention module outputs and are respectively input into the hierarchical disentangling and mining module, and the enhanced and are passed into the adaptive fusion module as inputs.
[0053] First, establish a disentangled feature initialization model, and define the original input features of the 0th layer as follows: ; When , set the initial disentangled features as follows: ; Among them, , , is the sequence length, is the feature dimension.
[0054] In the residual disentangling stage, strip the captured complementary information layer by layer through the attention mechanism, and retain the modal residual detail features. The calculation process is as follows: ; ; Among them, and are the disentangled features of RGB and thermal infrared at the th layer respectively, is the query vector generated by the RGB features of the th layer, is the key vector generated by the RGB features of the th layer, is the value vector generated by the RGB features of the th layer, is the key vector generated by the thermal features of the th layer, is the value vector generated by the thermal features of the th layer, is the disentangling level index, is the RGB residual disentangling output feature of the th layer, For the layer thermal modal residual unwrapping output feature, in the above calculation, the query, key, and value vectors are defined as follows: ; ; ; where is the linear mapping parameter of the layer, and the dimension is: ; .
[0055] The attention function is defined using the scaled dot - product calculation method as follows: Att ; where, is the Softmax function.
[0056] The above unwrapping process is iteratively executed layers. After extracting the residual complementary details, it enters the multi - level aggregation stage. The attention outputs of each layer are cumulatively represented as cross - layer complementary information as follows: ; ; The final depth - enhanced feature output is as follows: ; ; All the hierarchical weight parameters in the above hierarchical unwrapping module are not shared, and the feature mapping parameters adopt the standard initialization strategy: Use He initialization, and Use Xavier initialization.
[0057] To achieve the final fusion of multi - modal enhanced features, an adaptive fusion module is further established. This module is connected to the output end of the hierarchical unwrapping and mining module.
[0058] The module includes a feature evaluation unit and a probability weighting unit. The former generates the modal fusion weights, and the latter realizes modal fusion based on the weights.
[0059] First, perform the feature evaluation operation as follows: ; where, is the enhanced RGB / thermal modal feature, is the global average pooling, is the global max pooling, is the channel splicing operation, is the fused feature description vector, is the multi-layer perceptron, is the hyperbolic tangent activation function, is the rectified linear unit, is the modal weight coefficient.
[0060] Finally, weighted fusion is performed based on the above fusion coefficients to obtain the final multi-modal enhanced feature: ; wherein, is the fusion weight coefficient of the RGB modality, is the fusion weight coefficient of the thermal modality, is the enhanced RGB feature, is the enhanced thermal modality feature, is the final fusion feature.
[0061] The above fusion feature can be used for subsequent multi-modal downstream tasks such as detection, segmentation or recognition.
[0062] This module can improve the detail sensitivity and task adaptability of the cross-modal system by introducing a structured channel perception mechanism to dynamically adjust the expression ratio of RGB and thermal modalities during the fusion process.
[0063] The engineering implementation of the entire technical chain includes: The cross-modal attention module is used for coarse-grained complementary information interaction; The hierarchical disentanglement and mining module is used for extracting deep residual complementary information; The adaptive fusion module dynamically calculates the fusion weights based on the global semantics to achieve the final fusion.
[0064] The system structure of the present invention is clear, the functional boundaries of the modules are clear, and the connection relationships between the modules can be realized through the neural network graph structure, and it has the ability of direct deployment.
[0065] Please refer to the appendix Figure 5 , the tracking and prediction module, which is connected to the adaptive fusion module and outputs the target position and scale information.
[0066] This embodiment provides a tracking and prediction module. The module further introduces a tracking and prediction module on the basis of the adaptive fusion module, which is used to output the target existence probability and accurate spatial position and scale parameters.
[0067] The tracking and prediction module includes: a classification branch, a regression branch and a decoding unit.
[0068] The physical structure of the module includes: a feature input end, a convolutional prediction branch part, and a bounding box decoding part. The connection relationship is: The fused features output by the adaptive fusion module are input into the tracking and prediction module; Among them, is the number of spatial sampling points of the fused features, is the channel dimension.
[0069] First, a spatial distribution modeling structure based on the fused features is established, and the response activation map related to the target position in the fused features is extracted through two-dimensional convolution operations: ; Among them, represents the classification branch convolutional layer, and the output represents the target existence probability prediction map at each spatial position, is the final fused feature.
[0070] Subsequently, a regression branch is constructed to estimate the relative position and scale parameters of the target bounding box: ; Among them, is the regression branch convolutional operation, and the output , each position corresponds to four-dimensional regression parameters , is the output feature of the regression branch.
[0071] The above is the relative center point offset, is the width and height scaling ratio.
[0072] In the target position prediction stage, a set of anchor templates is used as the basic framework.
[0073] Through the decoding unit, the relative parameters in are mapped to the true coordinate positions of the target, and the decoding process is as follows: ; ; Among them: is the corresponding anchor box parameter; represents the exponential transformation, which is used to improve the stability of scale regression.
[0074] Finally, combining the calculation formula with , outputs the target probability and spatial box parameters: Output Among them: Output is the final output of the tracking network; is the target existence probability for each position; is the Sigmoid function, used to limit the output value to the interval [0, 1]; (x, y, w, h) is the target prediction box after regression decoding.
[0075] The entire tracking and prediction process is described as follows: First, based on the fused features establish a spatial response prediction model; Subsequently, extract the classification score map and the regression parameter map through the dual-path convolution branches respectively; Finally, map the parameters to the actual target box position through the anchor box decoding function.
[0076] By introducing an end-to-end classification and regression structure, this module can achieve spatial localization and scale discrimination under fused features, improving the prediction accuracy and response robustness of the multi-modal tracking system.
[0077] Test example: Please refer to Appendix Figures 6 - 19 , to verify the effectiveness of the multi-scale decoupled neural network (MSDNet) proposed in the present invention in multi-modal target tracking tasks, system tests and evaluations are now carried out in combination with multiple publicly available RGB-thermal infrared multi-modal target tracking datasets.
[0078] I. Experimental settings In this embodiment, MSDNet uses ToMP as the basic tracker and adopts the first four convolutional blocks of ResNet-50 as the feature extraction backbone. During training, the same standard data augmentation strategy as ToMP is adopted, and the network parameters are initialized based on the pre-trained model of ToMP-50.
[0079] In the training stage, MSDNet is optimized in an end-to-end manner. The optimizer is stochastic gradient descent (SGD), jointly minimizing the classification loss and the regression loss. The initial learning rates of the backbone network, the predictor, and the IoU predictor are set to 1×10 -5 , where the learning rates of the attention module CA, the decoupling module HDM, and the alignment module AF are set to 2×10 -5 . The overall training lasts for 50 epochs. After the 20th epoch, the learning rate is decayed by 0.5 times every 10 epochs. The training platform is PyTorch, and both training and testing are completed on a single NVIDIA RTX 3090 graphics card.
[0080] II. Test datasets and evaluation metrics Dataset selection: GTOT: It contains 50 video sequences, a total of 7.8K frames, with a relatively low resolution, covering 7 types of tracking challenges.
[0081] RGBT234: It contains 234 video sequences, a total of 116.7K frames, covering 12 types of tracking challenges.
[0082] LasHeR: It contains 1224 video sequences, a total of 734.8K frames, including 32 types of targets and 19 challenge attributes.
[0083] VTUAV: It contains 500 high-resolution video sequences, a total of 1.7M frames, divided into short-term (VTUAV-ST) and long-term (VTUAV-LT) subsets, covering 13 attributes.
[0084] Evaluation metrics: Precision Rate (PR): The minimum distance between the center of the predicted bounding box and the centers of the RGB and thermal infrared groundtruths. If it is less than the threshold, it is judged as correct.
[0085] Success Rate (SR): If the maximum overlap rate between the predicted bounding box and the groundtruth box is greater than 0.5, it is considered a success.
[0086] The PR threshold for GTOT is set to 5 pixels, and the rest of the datasets are set to 20 pixels; the SR threshold for all datasets is 0.5.
[0087] Comparison methods: In the experiment, 17 current mainstream advanced RGBT tracking algorithms are selected for comparison, including but not limited to ADRNet, HMFT, CMD, QAT, ViPT, TBSI, MPT, BAT, LSAR, STMT, MMSTC, CAT++, QueryTrack, UnTrack, OneTrack, SDSTrack, SiamTFA, etc., to verify the superiority of the method of the present invention.
[0088] The specific meanings of the algorithms are as follows: ADRNet: An object tracking algorithm for adaptively learning attribute-driven representation; HMFT: A hierarchical multimodal fusion object tracking algorithm; CMD: An object tracking algorithm for cross-modal distillation learning; QAT: An object tracking algorithm for quality perception and weighted residuals; ViPT: A visual prompt multimodal object tracking algorithm; TBSI: An object tracking algorithm based on the interaction between the template and the search area; MPT: An object tracking algorithm for maximizing the peak sidelobe ratio; BAT: Bidirectional Adapter Multimodal Object Tracking Algorithm; LSAR: Online Learning Samples and Adaptive Recovery Object Tracking Algorithm; LSAR: Online Learning Samples and Adaptive Recovery Object Tracking Algorithm; STMT: Multimodal Object Tracking Algorithm Based on Spatiotemporal Attention Mechanism; MMSTC: Multimodal Object Tracking Algorithm Based on Spatiotemporal Context; CAT++: Object Tracking Algorithm Driven by Challenge Attributes; QueryTrack: Object Tracking Algorithm with Joint Modal Query Fusion; UnTrack: Object Tracking Algorithm Extended from Unimodal to Multimodal; OneTrack: Visual Multimodal Object Tracking Algorithm with a Unified Framework; SDSTrack: Multimodal Object Tracking Algorithm with Self-Distillation Symmetric Adapter Learning; SiamTFA: Object Tracking Algorithm with Three-Stream Feature Aggregation Based on Siamese Architecture.
[0089] III. Overall Performance Test Results The MSDNet proposed in this invention performs as follows on different datasets: Table 1: Data Results of the Performance of MSDNet on Different Datasets
[0090] Among them, the SR performance of MSDNet on the LasHeR dataset is only slightly lower than that of TBSI (the difference is 1.6%), but on the four datasets of GTOT, RGBT234, VTUAV-ST, and VTUAV-LT, it is significantly better than the existing best comparison methods. For example, it is 0.7% higher than OneTrack on RGBT234, 3.9% higher than SiamTFA on VTUAV-ST, and 7.3% higher than MMSTC on VTUAV-LT.
[0091] IV. Attribute-Level Performance Analysis To further verify the adaptability of the method, a fine-grained evaluation was conducted on 13 challenge attributes of the LasHeR and VTUAV-ST datasets. The results show that MSDNet is slightly insufficient under the attributes of FL (Fast Motion), OV (Occlusion), and FO (Object Disappearance), mainly because the model does not currently integrate an object re-detection mechanism. However, under most challenge conditions such as SV (Scale Variation), FM (Fast Motion), HI (Illumination Change), LI (Low Illumination), and CM (Similar Interference), this method shows higher robustness, is better than the comparison methods, and reflects its significant advantage in mining multimodal weak matching information.
[0092] V. Visual Analysis In an actual scenario, we conducted a comparative experiment by comparing the algorithm of this application with multiple traditional algorithms. We verified the effectiveness of this algorithm during actual use through the tracking results of the following 4 groups of extreme scenarios. Figures 8 - 19 Under the 4 groups of extreme scenarios shown (bike-5, car-72, excavator-001, pedestrian-164), the MSDNet proposed by the present invention all demonstrated extremely strong robustness. The specific result analysis is as follows: The first scenario: bike-5 (see the appendix Figure 8 - appendix Figure 10 , which is one of the video sequences in the entire large dataset. The target to be tracked is a bicycle, and the number 5 indicates that this is the fifth video sequence of a bicycle in the dataset (because there are many video sequences of bicycles in the entire dataset)): Only SiamTFA and MSDNet can maintain tracking when the target is completely occluded; The second scenario: car-72 (see the appendix Figure 11 - appendix Figure 13 , which is one of the video sequences in the entire large dataset. The target to be tracked is a car, and the number 72 indicates that this is the 72nd video sequence of a car in the dataset (because there are many video sequences of cars in the entire dataset)): In the face of a large number of similar interferences, only HMFT and MSDNet can maintain precise positioning; The third scenario: excavator-001 (see the appendix Figure 14 - appendix Figure 16 , which is one of the video sequences in the entire large dataset. The target to be tracked is an excavator, and the number 001 indicates that this is the first video sequence of an excavator in the dataset (because there are multiple video sequences of excavators in the entire dataset)): When the appearance of the object changes drastically, MSDNet performs the most stably; The fourth scenario: pedestrian-164 (see the appendix Figure 17 - appendix Figure 19 , which is one of the video sequences in the entire large dataset. The target to be tracked is a pedestrian, and the number 164 indicates that this is the 164th video sequence of a pedestrian in the dataset (because there are many video sequences of pedestrians in the entire dataset)): In a low-light scenario, other methods completely fail, and only MSDNet can stably track.
[0093] These results show that MSDNet has the ability to deeply perceive information between modalities and robustly model the appearance of the target, and is suitable for complex and dynamic target tracking tasks.
[0094] In the above four groups of extreme scenarios, we mainly annotate the tracking results through the target box code algorithms of different colors. Each color represents a different algorithm. The color bar in front of the algorithm is used to indicate that the target box with the same color in the figure is the output result of that algorithm. In short, the color of the color bar represents which algorithm, and the target box with the same color as the color bar in the figure represents the tracking result output by that algorithm. In this way, we can intuitively compare the tracking effects of different algorithms through the target boxes of different colors. As shown in the attached Figure 8 - Attachment Figure 19 shown, the result position of GT tracking is shown as a green square in the attached Figure 8 - Attachment Figure 19 shown, the result position of ADRNet (IJCV-2021) tracking is shown as a blue square in the attached Figure 8 - Attachment Figure 19 shown, the result position of SiamTFA (TITS-2025) tracking is shown as an orange square in the attached Figure 8 - Attachment Figure 19 shown, the result position of LSAR (TCSVT-2024) tracking is shown as a dark gray-green square in the attached Figure 8 - Attachment Figure 19 shown, the result position of HMFT (CVPR-2022) tracking is shown as a purple square in the attached Figure 8 - Attachment Figure 19 shown, the result position of MSDNet tracking is shown as a red square in the attached Figure 8 - Attachment Figure 19 shown, the result position of QAT (ACMMM-2023) tracking is shown as a pink square in the attached Figure 8 - Attachment Figure 19 shown as a pink square.
[0095] Working principle: The system first extracts the basic features of the RGB and thermal infrared modalities through a two-stream ResNet-50 network with parameter sharing to ensure the alignment of the two modalities in a unified feature space. Subsequently, the cross-modal attention module establishes a two-way interaction channel: the thermal modality uses its own features as query vectors to aggregate complementary information in the RGB features through scaled dot-product attention, while the RGB modality performs the reverse operation to extract the effective information of the thermal features, forming preliminary enhanced features.
[0096] After these enhanced features enter the hierarchical disentanglement and mining module, the utilized complementary information is gradually stripped through three-level residual calculations. At each level, a query vector is generated from the current modality features, the attention distribution with the key-value pairs of the other modality is calculated, and the attention features are differentially stripped from the original features to iteratively update the feature representation. The attention features output by each level are cumulatively summed to form a deep complementary information flow, which is finally superimposed on the initial enhanced features to generate optimized bimodal features.
[0097] The adaptive fusion module generates a feature description vector based on global pooling statistics (average pooling and max pooling), and calculates the dynamic weights through a two-layer MLP network with Tanh activation. After the non-negativity of the weights is constrained by ReLU, the bimodal features are probabilistically weighted and fused, enabling the network to automatically adjust the modality contribution according to scene characteristics (such as light intensity, thermal target saliency). Finally, the fused features are input into the tracking prediction head, and through a parallel classification and regression double-branch structure, the target presence probability map and bounding box parameters are predicted respectively. After coordinate decoding, the target position (x, y) and scale (w, h) in the physical space are output to complete the end-to-end tracking decision.
[0098] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A hierarchical disentanglement fusion network for modal shared information in RGB-T object tracking, characterized in that Including: A dual-stream feature extraction module, which uses a parameter-sharing ResNet-50 network to extract features of RGB images and thermal infrared images respectively; A cross-modal attention module, which is connected to the dual-stream feature extraction module to enhance the feature interaction between the RGB and thermal modalities through a bidirectional attention mechanism; A hierarchical disentanglement and mining module, which is connected to the cross-modal attention module to mine deep complementary information between modalities through multi-level residual calculations; An adaptive fusion module, which is connected to the hierarchical disentanglement and mining module to achieve multi-modal feature fusion based on dynamic weight allocation; A tracking and prediction module, which is connected to the adaptive fusion module to output target position and scale information.
2. The modality-shared information hierarchical disentanglement fusion network for RGB-T object tracking according to claim 1, wherein The cross-modal attention module includes: A thermal modality enhancement path, which uses thermal features as query vectors and RGB features as key-value pairs to calculate attention weights through Formula 1; An RGB modality enhancement path, which uses RGB features as query vectors and thermal features as key-value pairs to calculate attention weights through Formula 2.
3. The modality-shared information hierarchical disentanglement and fusion network for RGB-T object tracking according to claim 2, wherein Specifically, Formula 1 in the thermal modality enhancement path is: ; Among them, is a cross-modal attention calculation function, represents the Softmax function, is the query vector generated by the thermal modality features through linear projection, is the key vector generated by the modality features through linear projection, is the value vector generated by the modality features through linear projection; is the scaling factor; Specifically, Formula 2 in the RGB modality enhancement path is: ; Among them, is the query vector generated by linear projection of RGB modality features, is the key vector generated by linear projection of thermal modality features, is the value vector generated by linear projection of thermal modality features, is the cross-modal attention calculation function.
4. The modality-shared information hierarchical disentanglement and fusion network for RGB-T object tracking according to claim 1, wherein The hierarchical disentanglement and mining module performs the following operations: Iteratively calculate residual features, and layer by layer strip the utilized complementary information through Formulas 3 and 4; Accumulate multi-level attention features, and aggregate deep complementary information through Formulas 5 and 6.
5. The modality-shared information hierarchical disentanglement fusion network for RGB-T target tracking according to claim 4, wherein Specifically, Formulas 3 and 4 in the iterative calculation of residual features are: ; ; Among them, and are the disentangled features of RGB and thermal infrared in the th layer, is the query vector generated from the RGB features of the th layer, is the key vector generated from the RGB features of the th layer, is the value vector generated from the RGB features of the th layer, is the key vector generated from the thermal features of the th layer, is the value vector generated from the thermal features of the th layer, is the disentanglement layer index, is the RGB residual disentanglement output feature of the th layer, is the thermal modality residual disentanglement output feature of the th layer.
6. The modality-shared information hierarchical disentanglement and fusion network for RGB-T target tracking according to claim 4, wherein Specifically, Formulas 5 and 6 in the accumulation of multi-level attention features are: ; ; Among them, is the total number of unwrapping levels, is the query vector generated by the thermal feature of the th layer, is the cross-modal attention calculation function, is the cumulative complementary feature from RGB to the thermal modality, is the cumulative complementary feature from the thermal to the RGB modality.
7. The modality-shared information hierarchical disentanglement fusion network for RGB-T object tracking according to claim 1, wherein The adaptive fusion module includes: A feature evaluation unit, which calculates modality weights through Formula 7; A probability weighting unit, which realizes feature fusion through Formula 8.
8. The modality-shared information hierarchical disentanglement fusion network for RGB-T object tracking according to claim 7, wherein Specifically, Formula 7 is: ; Among them, is the enhanced RGB / thermal modality feature, is the global average pooling, is the global max pooling, is the channel concatenation operation, is the fused feature description vector, is the multi-layer perceptron, is the hyperbolic tangent activation function, is the rectified linear unit, is the modality weight coefficient; Specifically, Formula 8 is: ; Among them, is the fusion weight coefficient of the RGB modality, is the fusion weight coefficient of the thermal modality, is the enhanced RGB feature, is the enhanced thermal modality feature, is the final fusion feature.
9. The modality-shared information hierarchical disentanglement and fusion network for RGB-T object tracking according to claim 1, wherein The tracking and prediction module includes: A classification branch, which predicts the target presence probability map; A regression branch, which outputs target bounding box parameters; A decoding unit, which converts the regression parameters into (x, y, w, h) coordinates.
Citation Information
Patent Citations
Infrared and visible light image fused multispectral target detection method and system
CN113688806A
RGB-D image semantic segmentation method based on multi-modal feature fusion
CN114549439A
Real-time RGBT target tracking method based on multi-modal interaction and multi-stage optimization
CN115170605A
RGBT tracking method and system based on edge perception and cross-modal relation mining
CN117036406A
RGBT real-time tracking method and system based on multi-mode interactive fusion
CN119339198A
Cited By
Unmanned aerial vehicle positioning method, device and equipment based on sequence observation and medium
CN120831630A
Optical-infrared target detection method of cross-modal attention fusion mechanism
CN120913023A
An optical-infrared target detection method of cross-modal attention fusion mechanism
CN120913023B
Infrared and visible light feature fusion method based on space channel decoupling coupling
CN121214132A
An infrared and visible light feature fusion method based on spatial channel decoupling coupling
CN121214132B