RGBT tracking network based on pixel-level fusion and use method

By adopting a lightweight pixel-level fusion adapter and a task-oriented progressive learning framework in the RGBT tracking network, the problem of insufficient information utilization in the existing technology is solved, and the tracking performance in complex scenarios is improved.

CN119942152AActive Publication Date: 2025-05-06ANHUI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510035392.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Existing pixel-level fusion methods are difficult to mine and utilize complementary information between modes in complex tracking scenarios, resulting in a large performance gap.

Method used

A RGBT tracking network based on pixel-level fusion is designed, using lightweight pixel-level fusion adapter (PFA) and a task-oriented progressive learning framework, which improves the ability to utilize information between modes through multi-expert adaptive distillation and decoupling representation fine-tuning strategies.

Benefits of technology

It realizes the effective mining and utilization of complementary information between modes in complex tracking scenarios, improves tracking performance, and narrows the performance gap with feature-level fusion methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942152A_ABST
    Figure CN119942152A_ABST
Patent Text Reader

Abstract

The invention provides an RGBT tracking network based on pixel-level fusion and a use method, and the RGBT tracking network comprises a pixel-level fusion adapter: firstly, each mode is divided by a low-level feature extraction layer, and then an independent Vim block is fed in to encode a specific feature; next, tokens and channel connections are applied to merge the two modalities along different feature dimensions, and the two additional Vim blocks further encode the fused information. And finally, decoding the fused features into an image by using a convolutional layer with efficient local detail modeling capability. The invention provides a two-stage task-oriented progressive learning framework. A first stage, multi-expert adaptive distillation (MAD). The method aims to inherit excellent fusion capability from various image fusion models with different structures. And in the second stage, decoupling represents a fine tuning strategy (DRF), task-related information and task-unrelated information are clearly separated through rejection loss to improve fusion precision, and completeness of information decoupling is ensured through reconstruction loss, so that fusion robustness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular to an RGBT tracking network based on pixel-level fusion and a use method thereof. Background Art

[0002] As a fundamental task in the field of computer vision, visual tracking aims to locate the target object in subsequent frames based on its initial state. Due to the complementary advantages of visible light (RGB) and thermal infrared (TIR) ​​modalities, which can enhance the tracking robustness in complex scenes, visible light thermal infrared (RGBT) tracking has received great attention from researchers in recent years, and a large number of works have shown impressive tracking performance. Current research mainly explores various fusion strategies, which can be divided into three categories, including pixel-level fusion, feature-level fusion, and decision-level fusion.

[0003] The first type of pixel-level fusion, which performs straightforward fusion directly at the pixel level. Yang et al., Prompting for multi-modal tracking. In Proceedings of the 30th ACM International Conference on Multimedia, pages 3492–3500, 2022. Exploring hint learning in multi-modal pixel-level interactions. And Tang et al., Exploring fusion strategies for accurate RGBT visual object tracking. Information Fusion, page 101881, 2023. Introducing an existing image fusion network for RGBT tracking.

[0004] The second type of feature-level fusion, due to the powerful representation ability of deep neural networks, feature-level fusion schemes have been widely used in RGBT tracking and have achieved remarkable results. Liu et al., Quality-aware RGBT tracking via supervised reliability learning and weighted residual guidance. In Proceedings of the 31st ACM International Conference on Multimedia, pages 3129–3137, 6682023. A feature-weighted fusion architecture based on the importance or quality of each modality in the last layer of features is adopted to achieve robust fusion. Chen et al., Simplifying cross-modal interaction via modality-shared features for RGBT tracking. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 1573–1582, 2024. After the fourth module, the channel-shared information aggregation module based on cross-attention is integrated into the ViT backbone to aggregate channel-shared information.

[0005] The third type of decision-level fusion processes the tracking results of visible light and thermal infrared separately, and then fuses the decisions of the two to finally obtain a comprehensive tracking result. This method is relatively simple and usually relies on optimizing the independent tracking algorithm for each modality. For example: Tang et al., Exploring fusion strategies for 697accurate RGBT visual object tracking. Information Fusion, 698page 101881, 2023. Zhang et al., Multi-modal 747fusion for end-to-end RGB-T tracking. In Proceedings of the 748IEEE International Conference on Computer Vision Work-749shops, 2019.

[0006] Due to the limited feature representation capabilities of shallow networks, current pixel-level fusion methods find it difficult to mine and utilize task-related inter-modal complementary information in complex tracking scenarios, resulting in a significant performance gap compared to feature-level fusion methods. Summary of the invention

[0007] The technical problem to be solved by the present invention is how to use a lightweight model to mine and utilize task-related inter-modal complementary information in complex tracking scenarios.

[0008] The present invention solves the above technical problems through the following technical means:

[0009] The RGBT tracking network based on pixel-level fusion is characterized by including:

[0010] Pixel-level fusion adapter: It includes three Vim blocks and two convolutional layers. The two modalities, RGB and TIR, are separated by a low-level feature extraction layer and then fed into the first Vim block to encode specific features. Then token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information. Finally, the fused features are decoded into an image to obtain a fused image.

[0011] Task-oriented progressive learning framework: including a task-oriented multi-expert distillation framework and a decoupled representation fine-tuning framework; the multi-expert distillation framework is used for parameter initialization of the pixel-level fusion adapter;

[0012] The decoupled representation fine-tuning framework includes two auxiliary branches and a tracking branch. During training, the two auxiliary branches operate in parallel with the tracking branch, and the two auxiliary branches are used to independently encode the task-independent information from RGB and TIR to obtain the RGB modality features. TIR modal characteristics Tracking branch for fusion image I f Perform feature extraction to obtain fusion features And perform the task loss; then combine It performs exclusion loss and fusion reconstruction loss to decouple the task relevance and task independence of the two modal features; outputs target classification and localization; during testing, the trained tracking branch is used for target tracking.

[0013] The present invention designs a lightweight pixel-level fusion adapter (PFA). First, each modality (RGB and TIR) is divided by a low-level feature extraction layer and then fed into a separate Vim block to encode specific features. Next, token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information. Finally, a convolutional layer with efficient local detail modeling capability is used to decode the fused features into an image.

[0014] The present invention proposes a two-stage task-oriented progressive learning framework (TPL). In the first stage, multi-expert adaptive distillation (MAD) is used to inherit superior fusion capabilities from multiple image fusion models with different structures. MAD can adaptively adjust the importance weights of different experts to ensure that the best experts are selected for distillation in different scenarios. In the second stage, the decoupled representation fine-tuning strategy (DRF) is used to explicitly guide PFA and tracker to integrate and extract task-related information. DRF improves fusion accuracy by explicitly separating task-related and task-irrelevant information through exclusion loss, and ensures the completeness of information decoupling through reconstruction loss, thereby improving fusion robustness.

[0015] As an optimization scheme of the above scheme, it also includes a nearest neighbor dynamic template update module: it is used to compare the predicted target classification score with a preset threshold, and then determine whether the model needs to update the target template.

[0016] As an optimization scheme for the above scheme, the multi-expert distillation framework library includes a CNN-based image fusion model, a diffusion-based image fusion model, and a Mamba-based image fusion model; given an RGB modality image and a TIR modality image, use I PFA , I c , I d , I m They represent the fused images of PFA, CNN-based image fusion model, diffusion-based image fusion model, and Mamba-based image fusion model respectively; therefore, the process of multi-expert adaptive distillation is as follows:

[0017]

[0018] in represents the IOU predictor for a single-stream tracker, represents the weight of the i-th expert; by minimizing the distillation loss To optimize the pixel-level fusion adapter parameters.

[0019] As an optimization solution of the above solution, the two auxiliary branches include a feature extraction trunk and a decoder; the tracking branch includes a feature extraction trunk with the same structure as the auxiliary branch, and a tracking head.

[0020] As an optimization solution of the above solution, the rejection loss is defined as:

[0021]

[0022] Where (x) + represents max(0,x), and α is the margin parameter that controls the similarity threshold.

[0023] As an optimization of the above scheme, the fusion reconstruction operation combines the task-related and task-irrelevant features and puts them into the decoders of the two auxiliary branches to reconstruct the original modality image. This process is expressed as:

[0024]

[0025] Among them, D rgb and D tir Represents RGB and TIR mode encoders respectively, represents the mean square error function.

[0026] The present invention also provides a pixel-level fusion method using the above-mentioned RGBT tracking network based on pixel-level fusion, comprising the following steps:

[0027] Steps of pixel-level fusion feature extraction: RGB and TIR are taken as inputs of the pixel-level fusion adapter, divided by a low-level feature extraction layer respectively, and then fed into the first Vim block to encode specific features; token and channel connections are then applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information; finally, the fused features are decoded into an image using

[0028] Task-oriented progressive learning: including the steps of initializing pixel-level fusion adapter parameters and fine-tuning the disentangled representation;

[0029] The step of adjusting the pixel-level fusion adapter parameters is to use a task-oriented multi-expert distillation framework to initialize the parameters of the pixel-level fusion adapter;

[0030] The steps of decoupled representation fine-tuning are as follows: during training, the two auxiliary branches operate in parallel with the tracking branch, and the two auxiliary branches are used to independently encode the task-independent information from RGB and TIR to obtain the RGB modality features. TIR modal characteristics Tracking branch for fusion image I f Perform feature extraction to obtain fusion features And perform the task loss; then combine It performs exclusion loss and fusion reconstruction loss to decouple the task relevance and task independence of the two modal features; outputs target classification and localization; during testing, the trained tracking branch is used for target tracking.

[0031] As an optimization scheme for the above scheme, the expert library includes a CNN-based image fusion model, a diffusion-based image fusion model, and a Mamba-based image fusion model; given an RGB modality image and a TIR modality image, use I PFA , I c , Id , I m They represent the fused images of PFA, CNN-based image fusion model, diffusion-based image fusion model, and Mamba-based image fusion model respectively; therefore, the process of multi-expert adaptive distillation is as follows:

[0032]

[0033] in represents the IOU predictor for a single-stream tracker, represents the weight of the i-th expert; by minimizing the distillation loss To optimize the pixel-level fusion adapter parameters.

[0034] As an optimization solution of the above solution, the two auxiliary branches include a feature extraction trunk and a decoder; the tracking branch includes a feature extraction trunk with the same structure as the auxiliary branch, and a tracking head.

[0035] As an optimization solution of the above solution, the rejection loss is defined as:

[0036]

[0037] Where (x) + represents max(0,x), α is the residual parameter that controls the similarity threshold;

[0038] The fusion reconstruction operation combines the task-related and task-irrelevant features and puts them into the decoders of the two auxiliary branches to reconstruct the original modality image. This process is expressed as:

[0039]

[0040] Among them, D rgb and D tir Represents RGB and TIR mode encoders respectively, represents the mean square error function.

[0041] The advantages of the present invention are:

[0042] This paper designs a lightweight pixel-level fusion adapter (PFA) that adopts the Mamba architecture with global modeling capabilities and linear computational complexity to achieve effective pixel-level fusion between two modal images. First, each modality (RGB and TIR) is divided by a low-level feature extraction layer and then fed into a separate Vim block to encode specific features. In order to enhance cross-channel representation, a parameter-free channel swapping mechanism is introduced to solve the limitation of the channel dimension of the Vim model. Next, token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information. Finally, a convolutional layer with efficient local detail modeling capabilities is used to decode the fused features into a fused image.

[0043] The present invention proposes a two-stage task-oriented progressive learning framework (TPL). In the first stage, multi-expert adaptive distillation (MAD) is used to inherit superior fusion capabilities from multiple image fusion models with different structures. MAD can adaptively adjust the importance weights of different experts to ensure that the best experts are selected for distillation in different scenarios. In the second stage, the decoupled representation fine-tuning strategy (DRF) is used to explicitly guide PFA and tracker to integrate and extract task-related information. DRF improves fusion accuracy by explicitly separating task-related and task-irrelevant information through exclusion loss, and ensures the completeness of information decoupling through reconstruction loss, thereby improving fusion robustness.

[0044] In order to solve the problem that the appearance difference between the initial template and the search frame in the pixel-level fusion scheme changes significantly over time, which in turn limits the performance of the pixel-level fusion tracker, this paper proposes a method that combines additional dynamic templates, the nearest neighbor dynamic template update strategy (NDTU), to bridge the appearance difference. In the tracking inference phase, the dynamic template is appropriately updated by identifying reliable frames. And in order to improve efficiency, only the latest reliable frames in history are saved, thereby ensuring the timely update of the template during the tracking process while avoiding the accumulation of too much redundant data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is the framework of the RGBT tracking network based on pixel-level fusion in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] In order to reveal the potential of pixel-level fusion in RGBT tracking, this embodiment proposes a new task-driven RGBT tracking network based on pixel-level fusion, such as Figure 1 As shown, the network has an efficient and effective progressive learning framework named TPF. In order to balance the efficiency and ability of pixel-level fusion, this embodiment uses Mamba with linear computational complexity to build a lightweight pixel-level fusion adapter (PFA), which contains only 14.3KB parameters to ensure the efficiency of image fusion. Subsequently, a new task-driven progressive learning (TPL) framework is proposed, including two stages of multi-expert adaptive extraction and decoupled representation fine-tuning to improve the fusion ability of PFA. In the first stage, the multi-expert adaptive extraction strategy guides PFA to inherit the advantages of different state-of-the-art image fusion models in different tracking scenarios, laying a solid foundation for fusion. In the second stage, a decoupled representation fine-tuning strategy is proposed to decouple RGB and TIR features into task-related and task-independent components, realizing a pixel-level fusion model focusing on task-related features. Specifically, it uses two auxiliary branches to independently encode task-independent information from RGB and TIR, while the tracking branch extracts task-related features in combination with the modality. This process is guided by three key losses, including the task loss to ensure relevance, the rejection loss to separate task-independent information, and the reconstruction loss to maintain input integrity. Therefore, through explicit guidance, PFA and tracker can integrate task-relevant information while filtering out task-irrelevant information. It is worth noting that the rejection loss and reconstruction loss are only valid during training. In addition, the change in target appearance between the initial template and the search frame over time poses additional challenges for pixel-level fusion. To overcome this, this embodiment designs a simple and effective nearest neighbor dynamic template update strategy (NDTU) to select the closest reliable tracking result to the current search area as the dynamic template.

[0048] The network of this embodiment consists of three parts: a pixel-level fusion adapter (PFA), a task-oriented progressive learning framework (TPL), and a nearest neighbor dynamic template update strategy (NDTU).

[0049] 1. To achieve pixel-level fusion, this embodiment does not directly concatenate or sum the image matrices. Drawing inspiration from the research on visible light and infrared image fusion, this embodiment designs a lightweight pixel-level fusion adapter (PFA). Different from existing computationally intensive image fusion methods, this embodiment adopts the Mamba architecture with global modeling capabilities and linear computational complexity to achieve effective pixel-level fusion between two modality images, which is better suited for tracking tasks.

[0050] 2. In order to balance the efficiency and capability of pixel-level fusion and improve the fusion capability of PFA, this embodiment proposes a two-stage task-oriented progressive learning framework (TPL). Specifically: In the first stage, multi-expert adaptive distillation. PFA can provide effective image fusion due to its low parameter setting, but it also limits the fusion performance. Although previous methods have attempted to adopt refinement strategies to improve fusion capabilities, it is difficult to provide sufficient guidance in various tracking scenarios relying solely on the knowledge of a single expert model. To overcome this limitation, this embodiment proposes a method called multi-expert adaptive extraction (MAD), which aims to inherit superior fusion capabilities from multiple image fusion models with different structures. MAD can adaptively adjust the importance weights of different experts to ensure that the best experts are selected for distillation in different scenarios. In the second stage, decoupled representation fine-tuning (DRF). Since all expert models are still oriented to human vision, further fine-tuning of PFA is necessary to effectively adapt to tracking tasks. To this end, this embodiment proposes a new decoupled representation fine-tuning strategy (DRF) to explicitly guide PFA and trackers to integrate and extract task-related information. DRF improves fusion accuracy by explicitly separating task-related and task-irrelevant information through the rejection loss, and improves the robustness of fusion by ensuring supervision with complete modality information through the reconstruction loss.

[0051] 3. Since the appearance difference between the initial template and the search frame in the pixel-level fusion scheme changes significantly over time, this also limits the performance of the pixel-level fusion tracker. To this end, this embodiment introduces a simple but effective nearest neighbor dynamic template update strategy (NDTU) to bridge the appearance difference by combining additional dynamic templates.

[0052] The following is a detailed introduction to the three parts of the network in this embodiment:

[0053] 1. Pixel Fusion Adapter (PFA)

[0054] This embodiment Figure 1 The network details of PFA are shown in . PFA mainly consists of three Vim blocks with channel sizes of 8, 8, and 16, and two convolutional layers. First, each modality (RGB and TIR) is divided by a low-level feature extraction layer and then fed into a separate Vim block to encode specific features. In order to enhance cross-channel representation, this embodiment introduces a parameter-free channel exchange mechanism to solve the limitation of the channel dimension of the Vim model. Next, token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information. Finally, this embodiment uses a convolutional layer with efficient local detail modeling capabilities to decode the fused features into an image.

[0055] 2. Task-based Progressive Learning Framework (TPL)

[0056] In the first stage, multi-expert adaptive distillation (MAD). Specifically, this embodiment selects three different image fusion models, including a CNN-based model (Zhao et al., Cddfuse: Correlation-driven dual-branch feature decomposition for multi-modality image fusion. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 5906–5916, 2023), a diffusion-based model (Yi et al., Diff-if: Multi-modality image fusion via diffusion fusion model with fusion knowledge prior. Information Fusion, 110: 102450, 2024.2, 3, 4, 7, 1 830) and a Mamba-based model (Li et al., Mambadfuse: A mamba-based dual-phase model for multi-modality image fusion. arXiv preprint arXiv: 2404.08406, 2024.2, 4, 7, 1740) as three experts of PFA. Given an RGB modality image and a TIR modality image, this embodiment uses I PFA , O c , I d , I m They represent the fused images of PFA, CNN-based image fusion model, diffusion-based image fusion model, and Mamba-based image fusion model, respectively. Therefore, the process of MAD is as follows:

[0057]

[0058] in represents the IOU (Intersection over Union) predictor for a single stream tracker, represents the weight of the i-th expert. By minimizing the distillation loss To optimize the PFA parameters, it inherits the existing multi-expert fusion capabilities and ensures lower computational complexity.

[0059] The second stage: Decoupled representation fine-tuning strategy (DRF). Specifically, the decoupled representation fine-tuning framework includes two auxiliary branches and a tracking branch. This embodiment introduces two auxiliary branches, each of which includes a feature extraction trunk with the same structure as the tracking branch and a modality-specific decoder. In addition to the feature extraction trunk, the tracking branch also includes a tracking head. These auxiliary branches operate in parallel with the tracking branch, and the two auxiliary branches are used to independently encode the task-independent information from RGB and TIR to obtain RGB modality features. TIR modal characteristics Tracking branch for fusion image I f Perform feature extraction to obtain fusion features And perform the task loss; then combine The exclusion loss and fusion reconstruction loss are performed to decouple the task relevance and task irrelevance of the two modality features ( Contains task-related information of two modes. The trained tracking branch is used to track the target. In order to separate the task-related and task-irrelevant features in each channel, this embodiment uses task loss to optimize And introduce a rejection loss to and Push Away Rejection loss is defined as:

[0060]

[0061] Where (x) + represents max(0,x), and α is the margin parameter that controls the similarity threshold. The rejection loss prevents unnecessary overlap between these features during fine-tuning, allowing the model to focus on task-relevant information in each modality and minimize interference.

[0062] In order to ensure the integrity of feature decoupling, this embodiment combines task-related and task-irrelevant features and puts them into the corresponding modality-specific decoder to reconstruct the original modality image. This process can be expressed as:

[0063]

[0064] Among them, D rgb and D tir denote the RGB and TIR mode encoders respectively, and Represents the mean squared error function. The reconstruction loss ensures that each channel retains both task-related and task-irrelevant information during fine-tuning, achieves dynamic adjustment of these features, and minimizes the loss of key task details during fusion.

[0065] Finally, this embodiment defines the final loss of the fine-tuning stage as a combination of reconstruction loss, rejection loss, and task-specific loss, as follows:

[0066]

[0067] in represents the tracking task loss. The coefficient λ is used to balance the contribution of each loss component. For simplicity, it is set to 1 in the study of this embodiment.

[0068] 3. Nearest Neighbor Dynamic Template Update Strategy (NDTU)

[0069] Since the appearance difference between the initial template and the search frame in the pixel-level fusion scheme changes significantly over time, this also limits the performance of the pixel-level fusion tracker. To this end, a simple but effective nearest neighbor dynamic template update strategy (NDTU) can be introduced to bridge the appearance difference by combining additional dynamic templates. In order to properly update the dynamic template during the tracking inference phase, this embodiment first identifies reliable frames, defined as frames whose predicted target classification score P exceeds the reliable threshold p. For all historically reliable frames, this embodiment only saves the latest reliable frame S new To improve efficiency. Every N frames, this embodiment introduces S new As the dynamic template of the current frame. In the training phase, this embodiment introduces dynamically updated templates sampled from intermediate frames as additional input. Benefiting from these templates are the fusion results of frames with different appearance differences from different moments, so the model can better capture the appearance changes of the target by modeling the global relationship between spatial and temporal dimensions.

[0070] This embodiment compares the performance of the algorithm of this application with the most advanced RGBT tracking method on the GTOT, RGBT210, RGBT234 and LasHeR datasets to verify the effectiveness of the proposed method (Table 1). This application selects the most commonly used evaluation indicators of each dataset for quantitative performance evaluation. The LasHeR dataset uses PR, NPR, and SR as evaluation indicators, and the other datasets use PR and SR as evaluation indicators. Among them, PR refers to the percentage of frames in which the distance between the center point of the tracking result and the center of the true bounding box is lower than the specified threshold. SR measures the ratio of frames in which the overlapping area of ​​the tracking result box and the true bounding box is higher than the threshold, and then calculates the SR metric by changing the threshold and calculating the area under the result curve (AUC). In order to eliminate the influence of image resolution and target size on the PR metric, PR is normalized to obtain the normalized accuracy rate NPR when evaluating the tracking performance on the LasHeR dataset. It can be seen from the results that the method proposed in this embodiment can be applied to the RGBT target tracking dataset, and the results are significantly improved compared with other methods.

[0071] Table 1 Comparison of method results

[0072]

[0073]

[0074] Table 1 Comparison of method results

[0075]

[0076]

[0077] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. RGBT tracking network based on pixel-level fusion, characterized by: include: Pixel-level fusion adapter: includes three Vim blocks and two convolutional layers; The two modalities, RGB and TIR, are separated by a low-level feature extraction layer and then fed into the first Vim block to encode specific features; Then token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information; finally, the fused features are decoded into an image using f ; Task-oriented progressive learning framework: including a task-oriented multi-expert distillation framework library and a decoupled representation fine-tuning module framework; the expert library is used for parameter adjustment initialization of the pixel-level fusion adapter; The decoupled representation fine-tuning module framework includes two auxiliary branches and one tracking branch; During training, the two auxiliary branches operate in parallel with the tracking branch, and the two auxiliary branches are used to independently encode the task-independent information from RGB and TIR to obtain the RGB modality features. TIR modal characteristics Tracking branch for fusion image I f Perform feature extraction to obtain fusion features And perform the task loss; then combine It performs exclusion loss and fusion reconstruction loss to decouple the task relevance and task independence of the two modal features; outputs target classification and localization; during testing, the trained tracking branch is used for target tracking.

2. The RGBT tracking network based on pixel-level fusion according to claim 1, characterized in that: It also includes a nearest neighbor dynamic template update module: it is used to compare the predicted target classification score with the preset threshold to determine whether the model needs to be updated.

3. The RGBT tracking network based on pixel-level fusion according to claim 1, characterized in that: The multi-expert distillation framework includes a CNN-based image fusion model, a diffusion-based image fusion model, and a Mamba-based image fusion model. Given an RGB modality image and a TIR modality image, use I PFA , I c , I d , I m They represent the fused images of PFA, CNN-based image fusion model, diffusion-based image fusion model, and Mamba-based image fusion model respectively; therefore, the process of multi-expert adaptive distillation is as follows: in represents the IOU predictor for a single-stream tracker, represents the weight of the i-th expert; by minimizing the distillation loss To optimize the pixel-level fusion adapter parameters.

4. The RGBT tracking network based on pixel-level fusion according to any one of claims 1 to 3, characterized in that: The two auxiliary branches include a feature extraction trunk and a decoder; the tracking branch includes a feature extraction trunk with the same structure as the auxiliary branches, and a tracking head.

5. The RGBT tracking network based on pixel-level fusion according to any one of claims 1 to 3, characterized in that: The rejection loss is defined as: Where (x) + represents max(0,x), and α is the margin parameter that controls the similarity threshold.

6. The RGBT tracking network based on pixel-level fusion according to any one of claims 1 to 3, characterized in that: The fusion reconstruction operation combines the task-related and task-irrelevant features and puts them into the decoders of the two auxiliary branches to reconstruct the original modality image. This process is expressed as: Among them, D rgb and D tir Represents RGB and TIR mode encoders respectively, represents the mean square error function.

7. A pixel-level fusion method using the RGBT tracking network based on pixel-level fusion according to any one of claims 1 to 6, characterized in that: The following steps are involved: Steps of pixel-level fusion feature extraction: RGB and TIR modalities are used as inputs of the pixel-level fusion adapter, divided by a low-level feature extraction layer respectively, and then fed into the first Vim block to encode specific features; Then token and channel connections are applied to merge the two modalities along different feature dimensions, and two additional Vim blocks further encode the fused information; finally, the fused features are decoded into an image using f ; Task-oriented progressive learning: including the steps of initializing pixel-level fusion adapter parameters and fine-tuning the disentangled representation; The step of adjusting the pixel-level fusion adapter parameters is to use a task-oriented multi-expert distillation framework to initialize the parameters of the pixel-level fusion adapter; The steps of decoupled representation fine-tuning are as follows: during training, the two auxiliary branches operate in parallel with the tracking branch, and the two auxiliary branches are used to independently encode the task-independent information from RGB and TIR to obtain the RGB modality features. TIR modal characteristics Tracking branch for fusion image I f Perform feature extraction to obtain fusion features And perform the task loss; then combine It performs exclusion loss and fusion reconstruction loss to decouple the task relevance and task independence of the two modal features; outputs target classification and localization; during testing, the trained tracking branch is used for target tracking.

8. The pixel-level fusion method according to claim 7, characterized in that: The multi-expert distillation framework includes a CNN-based image fusion model, a diffusion-based image fusion model, and a Mamba-based image fusion model. Given an RGB modality image and a TIR modality image, use I PFA , I c , I d , I m They represent the fused images of PFA, CNN-based image fusion model, diffusion-based image fusion model, and Mamba-based image fusion model respectively; therefore, the process of multi-expert adaptive distillation is as follows: in represents the IOU predictor for a single-stream tracker, represents the weight of the i-th expert; by minimizing the distillation loss To optimize the pixel-level fusion adapter parameters.

9. The pixel-level fusion method according to claim 7 or 8, characterized in that: The two auxiliary branches include a feature extraction trunk and a decoder; the tracking branch includes a feature extraction trunk with the same structure as the auxiliary branches, and a tracking head.

10. The pixel-level fusion method according to claim 7 or 8, characterized in that: The rejection loss is defined as: Where (x) + represents max(0,x), α is the residual parameter that controls the similarity threshold; The fusion reconstruction operation combines the task-related and task-irrelevant features and puts them into the decoders of the two auxiliary branches to reconstruct the original modality image. This process is expressed as: Among them, D rgb and D tir Represents RGB and TIR mode encoders respectively, represents the mean square error function.

Citation Information

Patent Citations

  • Cross-modal visual tracking method and device based on adaptive convolution

    CN114445462A

  • RGBT tracking method and system based on progressive fusion Transform and dynamic guidance learning

    CN116523956A

  • Multi-modal target tracking method and system based on coupling knowledge distillation

    CN118887255A

  • RGBT target tracking method based on multilayer global feature fusion and mapping template updating

    CN119169400A