RGB-T target tracking method based on multi-feature adaptive fusion and primary and secondary dynamic recovery
The RGB-T tracking method addresses instability and slow recovery in extreme lighting or occlusion scenarios by employing multi-feature adaptive fusion and dynamic recovery, ensuring robust and accurate tracking through adaptive modality switching and feature integration.
Patent Information
- Application Number
- CN202510811873.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing RGB-T tracking technology has problems such as inflexible modal fusion, insufficient resilience and incomplete feature expression, especially in extreme lighting or occlusion conditions.
Adaptive adjustment of modal features and stable tracking of targets is achieved through multi-level feature extraction, feature response fusion model and dynamic recovery mechanism.
It improves the tracking and discrimination ability in complex environments, enhances the complementarity of modal information, improves the recovery accuracy of targets after occlusion, and maintains the stability and continuity of target positioning.
Smart Images

Figure CN120318644A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to an RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery. Background Art
[0002] RGB-T tracking, as a hot research direction in multi-modal visual fusion in recent years, has shown great application potential in fields such as autonomous driving and intelligent monitoring. Its core advantage lies in combining the complementary characteristics of RGB images and thermal infrared images, and still having a certain degree of robustness under extreme lighting or occlusion conditions. However, there are still obvious deficiencies in the existing technology in terms of fusion methods and modal utilization mechanisms, and further breakthroughs are urgently needed.
[0003] On the one hand, most existing RGB-T tracking methods still adopt fixed weight or hierarchical splicing in the fusion strategy, and it is difficult to adapt to the importance changes of modal features in different scenarios. The modal quality fluctuates dynamically in actual tracking, especially in scenes with strong light, occlusion, or thermal infrared overexposure, and a single modality is prone to failure. This static fusion method not only lacks pertinence but also easily introduces redundant or misleading information, ultimately affecting tracking stability.
[0004] On the other hand, most existing methods rely on short-term tracking structures of specific modalities, and the recovery ability after tracking failure highly depends on global re-detection. Although this scheme is effective in some cases, when the target is occluded for a long time or the modality seriously degenerates, the recovery speed is slow and the accuracy is poor, making it difficult to meet the requirements of continuous tracking. Especially in decision-level fusion methods, the switching between the primary and secondary modalities often relies on fixed thresholds set manually, lacking a dynamic adaptation mechanism to cope with sudden changes in modal availability.
[0005] In addition, existing technologies generally focus on the modeling of a single feature type. Most methods only use single-layer features output by convolutional networks, ignoring the complementary relationship between multi-layer features, and rarely combine deep semantic features with manual features such as low-level edge textures. This limitation in feature expression leads to a significant decline in the discriminative ability of the model when the background is complex or the target deformation is severe, and it is extremely easy to produce drift.
[0006] Generally speaking, the current RGB-T tracking technology still has obvious shortcomings in terms of the flexibility of feature fusion, the ability of modal dynamic switching, and long-term robustness. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the present invention provides an RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery, which solves the problems of inflexible modal fusion, insufficient recovery ability, and incomplete feature expression in existing RGB-T tracking.
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: An RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery, comprising the following steps: S1. Perform multi-level feature extraction on the input RGB image and thermal imaging image respectively to obtain the feature sets of each modality; S2. Construct a multi-feature response fusion model based on the feature sets, and obtain the fusion response map and feature weight parameters through joint optimization calculation; S3. Calculate the tracking reliability evaluation value of the current frame according to the feature weight parameters; S4. When the tracking reliability evaluation value meets the preset conditions, perform short-term tracking optimization to determine the target position, and when it does not meet, perform long-term tracking recovery to re-capture the target; S5. Update the tracking model parameters according to the finally determined target position and output the tracking result.
[0009] Preferably, the step S2 includes: S21. Generate response maps for M features of each modality respectively, where M≥2; S22. Establish an optimization function with the goal of maximizing the peak significance of the response map and minimizing the response fluctuation; S23. Solve the feature weight parameters by using an optimization strategy combining particle swarm initialization and nonlinear programming.
[0010] Preferably, the optimization function is: ; The constraint conditions include: ; ; ; ; Among them, is the RGB modality feature weight vector, is the thermal imaging modality feature weight vector, is the multi-feature fusion response map, is the size of the response map, is the average value of the fusion response map, is the N×N all-ones matrix, is the th feature response map of the RGB modality, is the th feature response map of the thermal imaging modality, is the number of feature types, is the weight value of the th feature of the RGB modality, is the weight value of the th feature in the thermal imaging modality, represents the response value at the th row and th column in the fusion response map, represents the difference between the response value of each pixel in the fusion response map and the overall average response value, represents finding a set of optimal weights by optimizing the and weighted coefficients of the modality, so that the fused response map achieves the maximum discriminative ability in the current frame, is the maximum response value in the entire fusion response map.
[0011] Preferably, the tracking reliability evaluation value in step S3 is calculated by the following formula: ; where and are obtained by calculating through the HaarPSI algorithm, is the RGB modality template similarity, is the thermal imaging modality template similarity, is the joint modality confidence.
[0012] Preferably, the short-term tracking optimization in step S4 includes: generating a set of candidate regions associated with the historical trajectory; screening the optimal candidate through spatial constraint and appearance similarity calculation; refining the target bounding box based on the screening result.
[0013] Preferably, the screening of the optimal candidate adopts: ; where is the overlap ratio, is the spatial distance, is the area of the candidate box, is the confidence of the candidate box, is the maximum distance, is the appearance similarity, is the area of the reference target.
[0014] Preferably, the long-term tracking recovery in step S4 includes: generating an initial candidate box through the YOLOv4-tiny detector; selecting the top K = 5 high-confidence candidates for multi-modal response verification; selecting the candidate with the largest PSR response and exceeding the dynamic threshold as the recovery target, where the dynamic threshold takes the last valid PSR value before tracking failure.
[0015] The present invention provides an RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery, which has the following beneficial effects: 1. By adopting a multi-feature response adaptive fusion model, the present invention can dynamically identify the effectiveness of each modal feature in different scenarios, automatically adjust the fusion strategy, ensure the maximization of feature complementarity, and different from the existing fixed fusion ratio or simple splicing method, this solution breaks through the limitation of the coexistence of modal information redundancy and interference, and significantly improves the tracking and discrimination ability in complex environments.
[0016] 2. The present invention introduces a primary-secondary dynamic selection and recovery mechanism, enabling the tracking system to have the ability to autonomously adjust the modal priority and recovery path, and can re-capture the target under thermal infrared interference or extreme RGB illumination without human intervention. Compared with the traditional method relying on static modal switching strategies, this mechanism significantly improves the accuracy of target recovery after occlusion and effectively copes with the uncertainty risk brought by modal degradation.
[0017] 3. By using the combined response of the output of the deep convolutional network and the handcrafted features, the present invention not only solves the problem of insufficient shallow feature expression, but also retains the complementary advantages of high-level semantics and low-level details. Existing technologies often drift easily due to single feature expression when the target size changes or the background is cluttered, while this technical solution maintains the stability and continuity of target positioning through multi-feature fusion and response guidance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the flowchart of the method of the present invention; Figure 2 is the overall flowchart of MFDSR of the present invention; Figure 3 is the performance evaluation on the GTOT dataset of the present invention Figure 1 ; Figure 4 is the performance evaluation on the GTOT dataset of the present invention Figure 2 ; Figure 5 is the performance evaluation on the RGBT234 dataset of the present invention Figure 1 ; Figure 6 is the performance evaluation on the RGBT234 dataset of the present invention Figure 2 ; Figure 7 is the performance evaluation on the LasHeR dataset of the present invention Figure 1 ; Figure 8 is the performance evaluation on the LasHeR dataset of the present invention Figure 2 ; Figure 9 Performance evaluation on the VTUAV-ST dataset of the present invention Figure 1 ; Figure 10 Performance evaluation on the VTUAV-ST dataset of the present invention Figure 2 ; Figure 11 Performance evaluation on the VTUAV-LT dataset of the present invention Figure 1 ; Figure 12 Performance evaluation on the VTUAV-LT dataset of the present invention Figure 2 ; Figure 13 Visualization comparison graph of the tracking results of the present invention; Figure 14 Magnified comparison graph of 251 frames in the visualization comparison graph of the tracking results of the present invention; Figure 15 Magnified comparison graph of 1751 frames in the visualization comparison graph of the tracking results of the present invention; Figure 16 Magnified comparison graph of 5671 frames in the visualization comparison graph of the tracking results of the present invention; Figure 17 Magnified comparison graph of 451 frames in the visualization comparison graph of the tracking results of the present invention; Figure 18 Magnified comparison graph of 601 frames in the visualization comparison graph of the tracking results of the present invention; Figure 19 Magnified comparison graph of 4851 frames in the visualization comparison graph of the tracking results of the present invention. Detailed implementation manners
[0019] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Please refer to the attached Figure 1 - attached Figure 2 , the embodiments of the present invention provide an RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery, including the following steps: S1. Perform multi-level feature extraction on the input RGB image and thermal imaging image respectively to obtain the feature sets of each modality; Multi-level feature extraction includes a bimodal parallel processing architecture. For the input RGB image, first, the convolutional feature map of the 3rd layer (conv3) and the fully-connected feature map of the 14th layer (fc6) are extracted through the VGG-M network consisting of 13 convolutional layers and 3 fully-connected layers. At the same time, the Histogram of Oriented Gradients (HOG) features and Color Names (CN) features are calculated. The same operations are performed on the thermal imaging image synchronously to form a mixed feature set containing four types of features.
[0021] The specific implementation of VGG-M network feature extraction is as follows: The input image is normalized to a size of 125×125 pixels. After being processed by the convolutional layer group, the 3rd layer outputs a 28×28 spatial feature map with 64 channels, which can be mathematically expressed as: ; where, is the three-dimensional feature tensor output by the 3rd layer of VGG-M in the RGB modality, is the three-dimensional feature tensor output by the 3rd layer of VGG-M in the thermal imaging modality; The 14th layer outputs a 4096-dimensional fully-connected feature vector: ; where, is the fully-connected feature vector output by the 14th layer of VGG-M in the RGB modality, is the fully-connected feature vector output by the 14th layer of VGG-M in the thermal imaging modality; For HOG feature calculation, the cell size is set to 8×8 pixels, and the block size is 16×16 pixels (including 2×2 cells), generating a 31-dimensional feature descriptor: ; where, is the 31-dimensional descriptor of HOG features in the RGB modality, is the 31-dimensional descriptor of HOG features in the thermal imaging modality; For CN features, color namespace mapping is adopted to output an 11-dimensional color attribute histogram: ; where, is the 11-dimensional color histogram of CN features in the RGB modality, is the 11-dimensional color histogram of CN features in the thermal imaging modality; The construction of the feature set satisfies: ; ; where the superscripts 1-4 correspond to the VGG-M conv3 layer, fc6 layer, HOG, and CN features respectively, and the number of features M = 4. is the RGB modality feature set, is a thermal imaging modality feature set. The size of each feature map is based on the width of the target bounding box. Dynamic calculation: ; The logical connection between the feature extraction process and the fusion model is: HOG features provide edge structure information, CN features encode color distribution characteristics, VGG-M deep features capture semantic information, and shallow features retain spatial details. The four types of features form complementary representations through adaptive weighting, which is the subsequent optimization function: ; A differentiated feature substrate is provided, wherein: is an N×N matrix of all 1s, is the mean of the fused response graph, is the RGB modality feature weight vector, is the thermal imaging modality feature weight vector, is the multi-feature fusion response map, is the response graph size, is the average value of the fused response map, To represent the fusion response graph Row, No. The response value at the column, To represent the difference between the response value of each pixel in the fusion response map and the overall average response value, Indicates that through optimization and The weighted coefficients of the modalities are used to find a set of optimal weights so that the fused response graph can achieve the maximum discrimination ability in the current frame. It is the maximum response value in the entire fusion response graph.
[0022] In the embodiment, when processing a 640×480 resolution input image, the following operations are performed: intercepting 125×125 pixel RGB and thermal imaging dual-channel image blocks; extracting 64×28×28 feature tensors of the conv3 layer and 4096-dimensional feature vectors of the fc6 layer through the VGG-M network; calculating 31-dimensional HOG features and 11-dimensional CN features; and inputting the four types of features into the optimization model for weight distribution. This implementation method can realize the effective extraction of multi-scale features and provide spatial-semantic complementary feature representation for subsequent adaptive fusion.
[0023] S2, constructing a multi-feature response fusion model based on the feature set, and obtaining a fusion response graph and feature weight parameters through joint optimization calculation; The construction of the multi-feature response fusion model includes the following technical processes. First, the response graphs are generated for the M=4 features of each modality, where the RGB modality response atlas is , the thermal imaging modality response map set is . Response map size Dynamically calculated according to the width of the target bounding box : , where w is the relative width of the target normalized to the unit interval, represents the floor operation.
[0024] Establish an optimization function with the goal of maximizing the peak sidelobe ratio (PSR): ; The constraint conditions include: ; ; ; ; Among them, is the RGB modality feature weight vector, is the thermal imaging modality feature weight vector, is the multi-feature fusion response map, is the response map size, is the average value of the fusion response map, is the th feature response map of the RGB modality, is the th feature response map of the thermal imaging modality, is the number of feature types, is the th feature weight value of the RGB modality, is the th feature weight value of the thermal imaging modality, is an N×N all-ones matrix, satisfying , represents the response value at the th row and th column of the fusion response map, is the difference between each pixel response value and the overall average response value of the fusion response map, represents finding a set of optimal weights by optimizing the and modality weighted coefficients, so that the fused response map has the maximum discriminative ability in the current frame, is the maximum response value in the entire fusion response map.
[0025] Adopt an optimization strategy combining particle swarm initialization and interior point method, and the specific implementation is: Particle swarm initialization: Generate 60 particles, and each particle contains The velocity update formula for the dimensional parameters (M RGB weights + M TIR weights) is as follows: ; Among them, , the optimal solution is retained after 20 iterations , is the th iteration of the particle in the dimension velocity, is a uniformly distributed random number, is the particle in the dimension historical optimal position, is the th iteration of the particle in the dimension current position, is the global optimal position of the population in dimension dd, represents the th particle in the dimension, the th generation of velocity value; Interior point method optimization: Take as the initial point, set the initial value of the barrier factor , and the update strategy is: ; Among them, is the th iteration of the barrier factor, is the th iteration of the barrier factor; When the change in the objective function between adjacent iterations is less than , terminate.
[0026] In the embodiment, when processing a 125×125 pixel ROI, the following operations are performed: Perform bicubic interpolation on the 28×28 response map output by the 3rd layer of VGG-M to size; Calculate the weighted sum of the four types of feature response maps; After obtaining the initial weights through particle swarm optimization, the final output feature weight parameters are iteratively optimized by the interior point method. This implementation method can achieve the adaptive allocation of feature weights and improve the peak sidelobe ratio of the fusion response map.
[0027] S3. Calculate the tracking reliability evaluation value of the current frame according to the feature weight parameters; The calculation of the tracking reliability evaluation value is achieved through the following technical process. First, obtain the RGB module block and the thermal imaging module block of the current frame, and compare them with the high-confidence reference module blocks , Calculate the similarity. Use the HaarPSI algorithm to calculate the similarity: ; HaarPSI ; Among them, the HaarPSI algorithm calculates the local structural similarity through three-level wavelet decomposition. The specific implementation is as follows: Perform 3-layer Haar wavelet decomposition on the input module Calculate the structural similarity (SSIM) and phase consistency (PC) of each layer Weighted average the similarity scores of each layer, and the weight coefficient is Subsequently, calculate the joint modal confidence by combining the feature weight parameters: ; In the formula, is the RGB modal feature weight vector, is the thermal imaging modal feature weight vector, and the denominator realizes weight normalization, and are obtained by calculating through the HaarPSI algorithm, is the RGB modal template similarity, is the thermal imaging modal template similarity, is the joint modal confidence.
[0028] In the embodiment, when dealing with the occlusion scenario, perform the following operations: Perform HaarPSI similarity calculation on the module extracted from the th frame; Combine the optimized weights and , and calculate the joint modal confidence . When the average confidence of 5 consecutive frames is lower than 0.35, trigger the recovery mechanism.
[0029] S4. When the tracking reliability evaluation value meets the preset conditions, perform short-term tracking optimization to determine the target position. When it does not meet the conditions, perform long-term tracking recovery to recapture the target; The implementation of the primary and secondary dynamic recovery mechanism includes the following technical processes. When the average tracking reliability evaluation value of 5 consecutive frames is lower than the set threshold, perform long-term tracking recovery; otherwise, perform short-term tracking optimization.
[0030] The specific implementation of short-term tracking optimization is as follows: Generate a set of candidate regions that are spatially associated with the historical trajectory Calculate the candidate confidence through the following formula: ; Among them, is the overlap rate, is the spatial distance, is the area of the candidate box, is the confidence of the candidate box, is the maximum distance, is the appearance similarity, is the area of the reference target.
[0031] Select the candidate box with the highest confidence for bounding box refinement.
[0032] The specific implementation of long-term tracking recovery is as follows: Generate initial candidate boxes through the YOLOv4-tiny detector; Select the top K candidates with high confidence (K = 5); Perform multi-modal response verification on the candidate regions; Select candidates with PSR responses exceeding the dynamic threshold for target recovery.
[0033] In the embodiment, when long-term tracking recovery is triggered: Use the YOLOv4-tiny detector to generate candidate boxes; Select the top 5 candidates; Use the last valid PSR value before tracking failure as the dynamic threshold during the verification process; Re-initialize the tracker for candidates that meet the threshold conditions.
[0034] S5. Update the tracking model parameters according to the finally determined target position and output the tracking result.
[0035] The model update and result output process is implemented through the following technical solutions. Based on the finally determined target position, the ECO correlation filter architecture is used to update the model parameters. The specific implementation includes: Multi-feature template update: Perform respectively on the RGB and thermal imaging modalities: ; Among them, is the updated feature template, is the VGG-M feature (layers 3 and 14) extracted from the current frame, is the learning rate parameter, is the smoothed feature representation of the previous frame; Correlation filter parameter update: Adopt the ECO standard update strategy: ; Among them, , is the numerator / denominator term, is the regularization coefficient, is the filter / weight response function for the th frame.
[0036] Weight parameter update: Preserve historical weight information: ; where is the feature weight obtained by optimizing this frame, is the forgetting factor, is the fusion weight vector for the current frame, indicating that in the th frame, the fusion weight values corresponding to different modalities (such as RGB, TIR, etc.) determine the proportion of multi-modal features in the current frame fusion process, is the fusion weight vector of the previous frame.
[0037] In the embodiment, when the target position is determined: Extract multi-modal features of the target area; Update the VGG-M network feature template; Adjust the histogram of oriented gradients and color feature weights; Output the target coordinates and bounding box parameters.
[0038] For the test case, please refer to Appendix Figure 3 - Appendix Figure 13 : Test environment configuration: Use a hardware platform with an Intel i9 processor (3.0 GHz), NVIDIA RTX-3090 GPU, and 32 GB of memory, and run it in the Matlab-2022a software environment. The detector uses the default parameters of YOLOv4-tiny, and the correlation filter is implemented based on the ECO architecture.
[0039] Test method: Verify the algorithm on five benchmark datasets: GTOT, RGBT234, LasHeR, VTUAV-ST, and VTUAV-LT.
[0040] Set model parameters: The number of multi-features M = 4 (VGG-M layer 3 / 14 + HOG + Color), where HOG is the histogram of oriented gradients feature and Color is the color feature; The number of particles = 60, the termination tolerance = 10 −6 , the maximum number of iterations = 20; The number of long-term recovery candidates K = 5.
[0041] Evaluation metrics: Precision Rate (PR); Success Rate (SR)
[0042] Calculation of performance evaluation metrics: Meaning of precision: The distance error between the center of the target bounding box of each frame's tracking result and the center of the true annotation box. When the center distance error is less than 20 pixels, it is considered a correct tracking.
[0043] The calculation method of precision is: The total number of frames with a center error less than 20 pixels divided by the total number of frames in the sequence.
[0044] Meaning of success rate: The overlap rate between the target bounding box of each frame's tracking result and the true annotation box. When the overlap rate is greater than 0.5, it is considered a correct tracking.
[0045] The calculation method of success rate is: The total number of frames with an overlap rate greater than 0.5 divided by the total number of frames in the sequence.
[0046] For test results, please refer to Appendix Figure 3 - Appendix Figure 12 : The following table summarizes the performance of the algorithm of the present invention on each benchmark dataset:
[0047] Result analysis: According to the calculation methods of precision and success rate, we can obtain the precision and success rate result values of each tracker. The test results on each dataset will be analyzed in detail below.
[0048] The GTOT dataset, as shown in Appendix Figure 3 and Appendix Figure 4 As shown. On the GTOT dataset, the precision rate (PR) of the MFDSR algorithm of the present invention is 0.911, and the success rate (SR) is 0.756. Compared with other methods, the performance of the method of the present application is optimal on the GTOT dataset. This verifies the complementary fusion effect of deep features and handcrafted features. The algorithms participating in the comparison include: MFDSR: Multi-Feature Adaptive Fusion and Primary-Supplementary Dynamic Recovery Object Tracking Algorithm; DMCNet: Dual-Gated Mutual-Condition Object Tracking Algorithm; HMFT: Hierarchical Multi-Modal Fusion Object Tracking Algorithm; JMMAC: Joint Modeling of Motion and Appearance Object Tracking Algorithm; M5L: Multi-Modal Multi-Marginal Metric Learning Object Tracking Algorithm; MPT: Maximize Peak Sidelobe Ratio Object Tracking Algorithm; LSAR: Online Learning Samples and Adaptive Recovery Object Tracking Algorithm.
[0049] The RGBT234 dataset, as shown in AppendixFigure 5 and attached Figure 6 As shown. The accuracy of the MFDSR method on the RGB-T 234 dataset is 0.837, and the success rate is 0.582, which is basically equivalent (with little difference) to the accuracy of 0.840 and the success rate of 0.592 of CAT++. In the field of object tracking, generally, a performance difference of more than 1.0% is considered to have a large performance gap. The algorithms participating in the comparison include: CAT++: An object tracking algorithm driven by challenge attributes; MFDSR: A multi-feature adaptive fusion and primary-secondary dynamic recovery object tracking algorithm; M5L: An object tracking algorithm based on multi-modal multi-margin metric learning; JTPMA: A multi-modal multi-task feature fusion object tracking algorithm; LSAR: An object tracking algorithm for online learning samples and adaptive recovery; HMFT: A hierarchical multi-modal fusion object tracking algorithm; MPT: An object tracking algorithm that maximizes the peak sidelobe ratio; DFAT: An adaptive decision-level fusion tracking algorithm.
[0050] The LasHeR dataset, as attached Figure 7 and attached Figure 8 As shown. The PR / SR performance of MFDSR on the LasHeR dataset (0.475 / 0.395) significantly exceeds other state-of-the-art methods (compared with other methods, both in terms of accuracy and success rate, it is higher than 1%). Especially compared with the second-ranked CAT++ (an object tracking algorithm driven by challenge attributes), MFDSR achieves performance advantages of 3.1% / 3.3% in terms of PR / SR respectively. In addition, the method of this application has achieved good performance in most challenges. It is particularly noteworthy that LasHeR contains a relatively large number of long time series. In extremely challenging scenarios, the dynamic selection and recovery mechanism can handle these situations involving long-term tracking well, with an SR of 0.395, significantly better than other methods. The algorithms participating in the comparison include: MFDSR: A multi-feature adaptive fusion and primary-secondary dynamic recovery object tracking algorithm (of the present invention); CAT++: An object tracking algorithm driven by challenge attributes; APFNet: An object tracking algorithm based on attribute progressive fusion; DMCNet: A dual-gated mutual-condition object tracking algorithm; JTPMA: A multi-modal multi-task feature fusion object tracking algorithm; LSAR: An object tracking algorithm for online learning samples and adaptive recovery; MPT: Target tracking algorithm that maximizes the peak sidelobe ratio; DFAT: Adaptive decision-level fusion tracking algorithm.
[0051] VTUAV datasets (VTUAV-ST and VTUAV-LT), VTUAV-ST is as shown in the appendix Figure 9 and the appendix Figure 10 shown, VTUAV-LT is as shown in the appendix Figure 11 and the appendix Figure 12 shown.
[0052] The PR / SR performance of MFDSR on both the short-term (PR / SR: 0.773 / 0.626) and long-term (PR / SR: 0.530 / 0.453) tracking sub-datasets of VTUAV is significantly better than other algorithms. Specifically: In the VTUAV-ST (short-term) dataset, the success rate of MFDSR is basically the same as that of HMFT, but the accuracy is 1.5% higher than that of HMFT.
[0053] In the VTUAV-LT (long-term) dataset, the success rate of MFDSR is 9% higher than that of LSAR, and the accuracy is 10.3% higher than that of LSAR. These results indicate that the proposed scheme has strong robustness in complex environments. With the help of the adaptive fusion model and the dynamic selection and recovery mechanism, MFDSR performs well in most short-term and long-term tracking challenges.
[0054] The algorithms participating in the comparison in the VTUAV-ST dataset include: MFDSR: Multi-feature adaptive fusion and primary-secondary dynamic recovery target tracking algorithm (the present invention); HMFT: Hierarchical multi-modal fusion target tracking algorithm; MMMPT: Multi-modal interactive information transfer target tracking algorithm; LSAR: Target tracking algorithm with online learning samples and adaptive recovery; mfDIMP: End-to-end fusion target tracking algorithm; MPT: Target tracking algorithm that maximizes the peak sidelobe ratio; ADRNet: Target tracking algorithm with adaptive learning of attribute-driven representation DAFNet: Deep adaptive fusion target tracking algorithm The algorithms participating in the comparison in the VTUAV-LT dataset include: MFDSR: Multi-feature adaptive fusion and primary-secondary dynamic recovery target tracking algorithm (the present invention); HMFT: Hierarchical multi-modal fusion target tracking algorithm; LSAR: Target tracking algorithm with online learning samples and adaptive recovery; mfDIMP: End-to-end Fusion Object Tracking Algorithm; MPT: Object Tracking Algorithm for Maximizing Peak Sidelobe Ratio; ADRNet: Object Tracking Algorithm for Adaptive Learning Attribute-driven Representation; DAFNet: Deep Adaptive Fusion Object Tracking Algorithm; FSRPN: Object Tracking Algorithm for the Fusion of Siamese Network and Attention Mechanism.
[0055] Verification in Typical Scenarios: In an actual scenario, we verified the effectiveness of the algorithm in actual use by conducting comparative experiments between the algorithm of this application and multiple traditional algorithms and through the tracking results of the following two typical scenarios. The specific tracking results are as follows: Scenario 1: Large-scale Variation and Occlusion Scenario (corresponding to the upper part of the appendix, the detailed comparison results are shown in the appendix Figure 13 - appendix Figure 14 - appendix Figure 16 as shown): Large-scale variation means that the appearance of the target being tracked changes significantly, such as a change in size. Or when the shooting perspective changes, the appearance of the target also changes with the shooting perspective.
[0056] Occlusion means that the target being tracked is occluded by other objects, resulting in the inability to view the relevant information of the target in the image.
[0057] For the first scenario, the tracking target is a car in the video. In the appendix Figures 14 - 16 the tracking target is marked by a green box (the green box refers to GT). This scenario mainly targets the situation where the object scale changes greatly, exceeds the field of view, is severely occluded, and the thermal modality availability is poor (the thermal modality is specifically shown in the lower part of the appendix Figure 14 - appendix Figure 16 as shown). The entire video sequence of the scenario has 8107 frames. For the 8107 frames, we mainly extracted three typical frames for analysis. The schematic explanations of the algorithm tracking marks are as follows: The result position tracked by ADRNet is marked by a pink box in the appendix Figure 14 - appendix Figure 16 as shown. The result position tracked by MPT is marked by an orange box in the appendix Figure 14 - appendix Figure 16 as shown. The result position tracked by LSAR is marked by a dark gray-green box in the appendix Figure 14 - appendix Figure 16 as shown. The result position tracked by HMFT is marked by a yellow box in the appendix Figure 14 - appendix Figure 16 as shown. The result position tracked by MFDSR is in the appendixFigure 14 - Appendix Figure 16 is marked by a red square as shown in the figure.
[0058] As shown in the appendix Figure 14 When tracking reaches the 251st frame, the car undergoes severe deformation. The MFDSR algorithm can fully mark the position of the car and can handle the appearance changes of the target well. This is because our MFDSR algorithm incorporates an object detection algorithm, so it can accurately evaluate the size of the target. The green square is the true annotation box, whose main purpose is to mark the target we want to track, facilitating the direct comparison of the tracking results of different algorithms with the standard results. For other algorithms, since the aspect ratio of the target boxes output by their scale estimation modules is constant, for example, after scaling 1x2, it can only become 2x4, 4x8 target boxes with the same aspect ratio. However, in the actual tracking scenario, the height and width of the target change in any proportion, so they cannot handle large-scale changes well. When tracking reaches the 1751st frame, in a short period of time before the 1751st frame, the car passes through behind the tree (the car is blocked by the tree during the passing process). When the car appears again, that is, at the 1751st frame, the algorithm of this application can accurately re-locate the car to be tracked as shown in the appendix Figure 15 shown. While the other algorithms all fail to track. When tracking reaches the 5671st frame, the algorithm of this application can also robustly track the target car, at the position of the red square as shown in the appendix Figure 16 shown, while the other algorithms all fail. Through the above comparison examples, it can be seen that the dynamic selection and recovery mechanism of this application can re-capture the target and achieve continuous tracking.
[0059] Extreme illumination and complete occlusion scenarios (corresponding to the lower half of the appendix Figure 13 , the detailed comparison results are as shown in the appendix Figure 17 - Appendix Figure 19 shown): The extreme illumination and complete occlusion scenarios refer to low illumination scenarios plus occlusion scenarios.
[0060] For the second scenario, the tracking target is a pedestrian in the video. The tracking target in the appendix Figures 14 - 16 is marked by a green square (the green square refers to GT). This scenario mainly targets the situation where the object is in a low illumination scenario plus an occlusion scenario and the RGB modality has extremely poor availability. The entire video sequence of the scenario has 4987 frames. For the 4987 frames, we mainly extracted three typical frames for analysis. The schematic explanations of the algorithm tracking marks are as follows: The result position tracked by ADRNet is marked by a pink square as shown in the appendix Figure 17 - Appendix Figure 19 shown, and the result position tracked by MPT is marked by a pink square as shown in the appendix Figure 17 - Appendix Figure 19It is marked as an orange square in the figure, and the result position tracked by LSAR is in the appendix Figure 17 - appendix Figure 19 It is marked as a dark grayish green square in the figure, and the result position tracked by HMFT is in the appendix Figure 17 - appendix Figure 19 It is marked as a yellow square in the figure, and the result position tracked by MFDSR is in the appendix Figure 17 - appendix Figure 19 It is marked as a red square in the figure.
[0061] As shown in the appendix Figure 17 As shown, when tracking to the 451st frame, the tracked pedestrian is blocked by a tree, and at the 601st frame, the tracked pedestrian appears again (as shown in the appendix Figure 18 As shown). At this time, it can be seen that all algorithms other than the MFDSR algorithm of the present application fail to track, while the algorithm of the present application tracks correctly. When continuing to track to the 4851st frame, the algorithm of the present application is still very accurately tracking the target, while all other algorithms fail, and the position of the target pedestrian cannot be accurately tracked. Through the above examples, it can be seen that the adaptive fusion mechanism of the algorithm of the present application can identify modal availability and optimize weight configuration, effectively maintaining tracking stability.
[0062] The appendix of the present application Figure 13 - appendix Figure 19 The algorithms participating in the comparison in the appendix specifically include: MFDSR: Multi-Feature Adaptive Fusion and Master-Slave Dynamic Recovery Target Tracking Algorithm; HMFT: Hierarchical Multi-Modal Fusion Target Tracking Algorithm; LSAR: Target Tracking Algorithm with Online Learning Samples and Adaptive Recovery; MPT: Target Tracking Algorithm to Maximize Peak Sidelobe Ratio; ADRNet: Target Tracking Algorithm with Adaptive Learning Attribute-Driven Representation; GT: True Annotation of Ground Target Bounding Box (Label).
[0063] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery, characterized in that It includes the following steps: S1. Perform multi-level feature extraction on the input RGB image and thermal imaging image respectively to obtain the feature sets of each modality; S2. Build a multi-feature response fusion model based on the feature sets, and obtain the fusion response map and feature weight parameters through joint optimization calculation; S3. Calculate the tracking reliability evaluation value of the current frame according to the feature weight parameters; S4. When the tracking reliability evaluation value meets the preset conditions, perform short-term tracking optimization to determine the target position. When it does not meet the conditions, perform long-term tracking recovery to re-capture the target; S5. Update the tracking model parameters according to the finally determined target position and output the tracking result.
2. The RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery according to claim 1, characterized in that The step S2 includes: S21. Generate response maps for M features of each modality, where M≥2; S22. Establish an optimization function with the goal of maximizing the peak significance of the response map and minimizing the response fluctuation; S23. Adopt an optimization strategy combining particle swarm initialization and nonlinear programming to solve the feature weight parameters.
3. The RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery according to claim 2, characterized in that The optimization function is: ; The constraint conditions include: ; ; ; ; Among them, is the RGB modal feature weight vector, is the thermal imaging modal feature weight vector, is the multi-feature fusion response map, is the size of the response map, is the average value of the fusion response map, is an N×N all-ones matrix, is the th feature response map of the RGB modality, is the th feature response map of the thermal imaging modality, is the number of feature types, is the weight value of the th feature of the RGB modality, is the weight value of the th feature of the thermal imaging modality, represents the response value at the th row and th column in the fusion response map, represents the difference between each pixel response value and the overall average response value in the fusion response map, represents finding a set of optimal weights by optimizing the weighted coefficients of and modalities, so that the fused response map achieves the maximum discriminative ability in the current frame, is the maximum response value in the entire fusion response map.
4. The RGB-T object tracking method based on multi-feature adaptive fusion and primary and secondary dynamic recovery according to claim 3, wherein, In the step S3, the tracking reliability evaluation value is calculated by the following formula: ; Among them, and are obtained by calculating with the HaarPSI algorithm. is the similarity of the RGB modality template. is the similarity of the thermal imaging modality template. is the combined modality confidence.
5. The RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery according to claim 1, wherein In the step S4, the short-term tracking optimization includes: Generate a set of candidate regions associated with the historical trajectory; Screen the optimal candidate through spatial constraint and appearance similarity calculation; Refine the target bounding box based on the screening result.
6. The RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery according to claim 5, wherein The screening of the optimal candidate adopts: ; Among them, is the overlap rate, is the spatial distance, is the area of the candidate box, is the confidence of the candidate box, is the maximum distance, is the appearance similarity, is the area of the reference target.
7. The RGB-T object tracking method based on multi-feature adaptive fusion and primary-secondary dynamic recovery according to claim 1, characterized in that In the step S4, the long-term tracking recovery includes: Generate an initial candidate box through the YOLOv4-tiny detector; Select the top K = 5 high-confidence candidates for multi-modal response verification; Select the candidate with the largest PSR response and exceeding the dynamic threshold as the recovery target, where the dynamic threshold takes the last valid PSR value before tracking failure.
Citation Information
Patent Citations
RGBT target tracking method based on twin network structure and anchor frame adaptive thought
CN116563343A
Complex scene single-target tracking method, device and system based on Steple algorithm and storage medium
CN118864524A