Air-ground infrared time-sensitive target tracking method based on multi-mode correlation filtering
By building a multi-mode correlation filter model, using spatiotemporal regularization and sparse response constraint training filters, the robust tracking of space-ground infrared time-sensitive targets is achieved, and the problem of observation model degradation in space-ground infrared time-sensitive target tracking is solved, and the robustness and reliability of tracking are improved.
Patent Information
- Application Number
- CN202510312467.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
In the tracking of infrared time-sensitive targets in open-ground, existing related filtered target tracking methods are susceptible to target motion, appearance changes and scene complexity, resulting in degradation of observation models and degradation of tracking performance. Especially in the case of thermal crossover, infrared targets are confused with background and are difficult to classify.
Build infrared single-mode, RGB-T multi-mode fusion and visible light-assisted infrared observation models, train related filters through spatiotemporal regularization and sparse target response constraints, and use modal reliability adaptive decision fusion to achieve multi-mode switching and complementarity to improve tracking performance.
Achieve robust open-ground infrared time-sensitive target tracking in complex scenarios, improving the robustness and reliability of tracking, being able to flexibly respond to target motion and scene changes, and improving tracking performance in thermal crossover scenarios.
Smart Images

Figure CN120259364A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual target tracking for aircraft, and particularly to a method for tracking air-to-ground infrared time-sensitive targets based on multi-mode correlation filtering. Background Art
[0002] As an important part of visual target tracking, the air-to-ground infrared time-sensitive target tracking technology has excellent characteristics such as all-weather operation, passive non-contact, and anti-electromagnetic interference, and is widely used in fields such as precision-guided weapons, military reconnaissance and defense, and unmanned cruising. The air-to-ground scenario has characteristics such as strong target mobility, fast scene dynamic changes, and large background interference. How to achieve fast, accurate, and robust tracking of air-to-ground time-sensitive targets under complex conditions such as cluttered backgrounds, motion changes, and thermal crossovers is still the key and difficult point of infrared visual target tracking technology. The correlation filtering target tracking method can well balance rapidity, accuracy, and robustness, and has advantages such as lightweight computing requirements and low power consumption, and has become the main research direction in the field of air-to-ground target tracking.
[0003] Correlation filtering target tracking belongs to the discriminant method, and its observation model uses a correlation filter to classify the target and the background. In the application of air-to-ground infrared time-sensitive target tracking, affected by factors such as the motion characteristics of the aircraft, target mobility, and scene complexity, the correlation filtering target tracking is prone to serious degradation of the observation model, resulting in a significant decline in tracking performance. The correlation filtering target tracking establishes an observation model by training a correlation filter, and the classification ability of the filter determines the performance of the observation model. To achieve robust and reliable air-to-ground infrared time-sensitive target tracking, it is necessary to reasonably set the objective function to train the correlation filter to construct a robust observation model.
[0004] However, in the application of air-to-ground infrared time-sensitive target tracking, due to factors such as target motion, appearance changes, and scene complexity, the observation model of the correlation filtering target tracking method is prone to serious degradation, resulting in tracking drift at best and losing the target at worst. In addition, when there is a thermal crossover phenomenon in the tracking scene, the infrared target and the background are confused, and it is difficult for the correlation filter to effectively classify the two, resulting in a serious decline in the infrared target tracking performance.
[0005] To solve the above problems, there is an urgent need to provide a new method for tracking air-to-ground infrared time-sensitive targets based on multi-mode correlation filtering, which can flexibly select the tracking mode according to the application scenario and robustly track the air-to-ground infrared time-sensitive targets. Summary of the Invention
[0006] The purpose of this application is to provide a method for tracking air-to-ground infrared time-sensitive targets based on multi-mode correlation filtering. The method for tracking air-to-ground infrared time-sensitive targets based on multi-mode correlation filtering can flexibly select the tracking mode according to the application scenario and robustly track the air-to-ground infrared time-sensitive targets.
[0007] To achieve the above object, the present application provides the following solutions:
[0008] In a first aspect, the present application provides a method for tracking time-sensitive air-ground infrared targets based on multi-modal correlation filtering. The method for tracking time-sensitive air-ground infrared targets based on multi-modal correlation filtering includes:
[0009] Construct an infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model according to the tracking scenario; the infrared single-modal observation model calculates the target response map using a displacement correlation filter, estimates the target size using a scale correlation filter; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model; the RGB-T multi-modal fusion observation model calculates the target responses of the visible light and infrared modalities respectively using a displacement correlation filter; calculates the tracking reliability of each modality using the target responses; adaptively makes a decision fusion to predict the target state according to the modality reliability; and jointly trains the displacement correlation filters of the visible light and infrared modalities according to spatio-temporal regularization and the consistency constraints of the sparse target responses of the visible light and infrared, and trains the scale correlation filter through an incrementally updated appearance model in the visible light image; the visible light-assisted infrared observation model calculates the target response of the infrared modality using a displacement correlation filter; determines the tracking reliability using the target response of the infrared modality; predicts the target state according to the tracking reliability; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model.
[0010] When the aircraft conducts time-sensitive air-ground infrared target tracking, determine the corresponding observation model according to the tracking scenario; the observation models include: an infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model;
[0011] Use the observation model corresponding to the tracking scenario to perform target tracking on the current video frame.
[0012] Optionally, the infrared single-modal observation model calculates the target response map using a displacement correlation filter and estimates the target size using a scale correlation filter, specifically including:
[0013] Extract the histogram of oriented gradients features, grayscale features, and motion features of the target search area in the infrared image;
[0014] Use the displacement correlation filter to estimate the target response map based on the histogram of oriented gradients features, grayscale features, and motion features;
[0015] Use the scale correlation filter to estimate the target size based on the histogram of oriented gradients features.
[0016] Optionally, the infrared single-modal observation model utilizes
[0017]
[0018] to train the displacement-related filter h 1t ;
[0019] where is the displacement-related filter of the d-th dimension in the t-th frame, is the vectorized feature map of the d-th dimension of the image patch used to train the displacement-related filter h 1t in the t-th frame, D1 is the displacement-related filtering dimension, is the displacement-related filter of the d-th dimension in the (t - 1)-th frame, is the vectorized ideal target response, is the set of real numbers, ★ is the correlation operation operator, ⊙ is the Hadamard product, is the spatial regularization constraint weight, λ1 is the coefficient of the spatial regularization constraint term, λ2 is the coefficient of the temporal regularization constraint term, λ3 is the coefficient of the sparse target response constraint term, T is the length of the vector of the vectorized feature map for each channel, is the Frobenius norm F, ||||1 is the vector l1 norm.
[0020] Optionally, determining the tracking reliability using the infrared modal target response specifically includes:
[0021] Using the formula θ 1t = APCE(R 1t ) max(R 1t ) to determine the tracking reliability θ 1t ;
[0022] where, APCE(R 1t ) is the average peak correlation energy ratio of the infrared modal target response R 1t , max(R 1t ) is the maximum value of the infrared modal target response R 1t .
[0023] Optionally, predicting the target state according to the tracking reliability specifically includes:
[0024] When θ 1t ≥ θ 1t-1 , using the infrared modal target response R 1t to predict the target position, and estimating the target size based on the histogram of oriented gradients features using the scale-related filter at the currently predicted target position; where, θ 1t-1 is the tracking reliability in the (t - 1)-th frame;
[0025] When θ1t < θ 1t-1 When, the RGB-T multi-modal fusion observation model is used to predict the target state.
[0026] According to the specific embodiments provided by the present application, the present application has the following technical effects:
[0027] The present application provides a method for tracking air-ground infrared time-sensitive targets based on multi-modal correlation filtering. An infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model are constructed according to the tracking scenario, that is, the infrared single-modal and RGB-T multi-modal tracking modes are flexibly switched according to the actual application scenario. In the present application, aiming at the problem that the observation model of correlation filtering target tracking is severely degraded due to factors such as target motion, appearance change, and scene complexity, the discrimination ability of the correlation filter is improved by spatio-temporal regularization and sparse target response constraint; aiming at the problem that it is difficult for infrared target tracking to effectively cope with the thermal cross-tracking scenario, the visible light modality is adaptively fused according to the modal reliability, and the infrared target tracking performance is improved by using the complementarity of the two modalities based on the consistency of the sparse target response; the present application uses visible light to assist infrared target tracking, which is beneficial to the reliable and stable tracking of air-ground infrared time-sensitive targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is a schematic flowchart of a method for tracking air-ground infrared time-sensitive targets based on multi-modal correlation filtering in an embodiment of the present application;
[0030] Figure 2 It is a schematic overall flowchart of the infrared single-modal target tracking mode;
[0031] Figure 3 It is a schematic overall flowchart of the RGB-T multi-modal fusion target tracking mode;
[0032] Figure 4 It is a schematic overall flowchart of the visible light-assisted infrared target tracking mode;
[0033] Figure 5 It is a schematic diagram of partial tracking results (parts (a), (b), and (c) of 5 respectively adopt the infrared single-modal target tracking mode, the RGB-T multi-modal fusion target tracking mode, and the visible light-assisted infrared target tracking mode). DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, in combination with the accompanying drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0035] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] In an exemplary embodiment, as Figure 1 shown, a method for tracking time-sensitive air-ground infrared targets based on multi-modal correlation filtering is provided. The method includes the following S101 to S103. Among them:
[0037] S101, construct an infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model according to the tracking scenario;
[0038] The infrared single-modal observation model uses a displacement correlation filter to calculate the target response map, and uses a scale correlation filter to estimate the target size; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model; according to the infrared image, extract the histogram of oriented gradient features (Histogram of Oriented Gradient, HOG), gray-scale features, and motion features of the target search area; use the displacement correlation filter to estimate the target response map based on the histogram of oriented gradient features, gray-scale features, and motion features; use the scale correlation filter to estimate the target size based on the histogram of oriented gradient features.
[0039] The RGB-T multi-modal fusion observation model uses a displacement correlation filter to calculate the target responses of the visible light and infrared modalities respectively; calculates the tracking reliability of each modality using the target responses; adaptively makes a decision fusion to predict the target state according to the modality reliability; and jointly trains the displacement correlation filters of the visible light and infrared modalities according to spatio-temporal regularization and the consistency constraints of the visible light and infrared sparse target responses, and trains the scale correlation filter through an incrementally updated appearance model in the visible light image; when estimating the target position in the visible light modality, use HOG features, color space (Color Names, CN), and gray-scale features, and the modality reliability is determined according to the maximum value of the target response and the average peak-to-correlation energy ratio (Average Peak-to-correlation Energy, APCE).
[0040] The visible light assisted infrared observation model calculates the infrared modal target response using a displacement correlation filter; determines the tracking reliability using the infrared modal target response; predicts the target state based on the tracking reliability; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model.
[0041] Specifically, as Figure 2 shown, based on the correlation filter to predict the target state, the infrared single-modal observation model first extracts the HOG features, gray-scale features, and motion features of the target search area in the infrared image, and calculates the target response map using the displacement correlation filter.
[0042] Let represent the vectorized feature of the target search area image block in the t-th frame in the infrared modality, be the displacement correlation filter trained in the (t - 1)-th frame. At this time, the target response is calculated as follows:
[0043]
[0044] In the formula, is the inverse discrete Fourier transform. The target position in the t-th frame is determined according to the position of the maximum value of R 1t . Subsequently, at the current target prediction position, the scale correlation filter is used to estimate the target size based on the HOG features.
[0045] In each frame of the infrared image, the displacement correlation filter is trained according to spatio-temporal regularization and sparse target response constraints, and the scale correlation filter is trained through an incrementally updated appearance model.
[0046] Let represent the vectorized feature map of the image block used for training the displacement correlation filter in the t-th frame in the infrared modality . The displacement correlation filter is trained by minimizing the following objective function:
[0047]
[0048] Among them, is the d-th dimensional displacement correlation filter in the t-th frame, is the d-th dimensional vectorized feature map of the image block used for training the displacement correlation filter h 1t in the t-th frame. D1 is the dimension of the displacement correlation filter, is the d-th dimensional displacement correlation filter in the (t - 1)-th frame, is the vectorized ideal target response, is the set of real numbers, ★ is the correlation operation operator, and ⊙ is the Hadamard product, is the weight of the spatial regularization constraint, λ1 is the coefficient of the spatial regularization constraint term, λ2 is the coefficient of the temporal regularization constraint term, λ3 is the coefficient of the sparse target response constraint term, and T is the length of the vector of the vectorized feature map for each channel. is the Frobenius norm F, and || ||1 is the l1 norm of the vector. The above objective function is solved by the alternating direction method of multipliers to obtain the displacement correlation filter h for predicting the target position in the (t + 1)-th frame. 1t .
[0049] In the RGB-T multi-modal fusion observation model, the displacement correlation filter is used to calculate the target responses of the visible light and infrared modalities respectively. Let represent the vectorized feature of the image patch in the target search area of the k-th modality (k = 1 represents the infrared modality, k = 2 represents the visible light modality) in the t-th frame. be the displacement correlation filter trained in the (t - 1)-th frame, and at this time the target response is calculated as follows:
[0050]
[0051] where is the inverse discrete Fourier transform; T is the vector length; and D k is the dimension.
[0052] The target state is predicted by adaptively fusing decisions based on the modality reliability. The visible light modality and the infrared modality are adaptively fused, and the final target response map in the t-th frame is calculated as follows:
[0053]
[0054] The target position in the t-th frame is determined according to the position of the maximum value of R 1t . Subsequently, at the current target prediction position, the scale correlation filter is used to estimate the target size in the visible light image.
[0055] As Figure 3 shown, for each frame, the two-modal displacement correlation filters are jointly trained according to the spatio-temporal regularization and the consistency constraint of the visible light and infrared sparse target responses, and the scale correlation filter is trained by the incrementally updated appearance model in the visible light image.
[0056] Let represent the vectorized feature map of the image patch for training the displacement correlation filter in the t-th frame of the k-th modality (k = 1 represents the infrared modality, k = 2 represents the visible light modality). The two-modal displacement correlation filters are jointly trained by minimizing the following objective function:
[0057]
[0058] In the formula, is the vectorized ideal target response, is the spatial regularization constraint weight, λ1 is the coefficient of the spatial regularization constraint term, λ2 is the coefficient of the temporal regularization constraint term, and λ3 is the coefficient of the sparse target response consistency constraint term. When the visible light modality input is zero, the visible light and infrared sparse target response consistency constraint term degenerates into the sparse target response constraint term, and the objective function of the RGB-T multi-modal fusion observation model is equivalent to the objective function of the infrared single-modal observation model. It can be seen from this that the method proposed in this application can realize the free switching between single-modal target tracking and RGB-T multi-modal target tracking.
[0059] In the visible light assisted infrared observation model, let represent the vectorized feature of the target search area image block in the infrared modality at the t-th frame,
[0060] is trained by the displacement-related filter trained in the (t - 1)-th frame. At this time, the target response is calculated as follows:
[0061]
[0062] Use the infrared modality target response to judge its tracking reliability. Let max(R 1t ) and APCE(R 1t ) represent the maximum value and APCE of R 1t respectively. At this time, the reliability θ 1t of the t-th frame in the infrared modality is calculated as follows:
[0063] When θ 1t ≥θ 1t-1 , it indicates that the current infrared modality tracking is stable and reliable; when θ 1t <θ 1t-1 , it indicates that the current infrared modality tracking reliability is low, and the visible light modality needs to be added for assisted tracking.
[0064] As Figure 4 shown, when θ 1t ≥θ 1t-1 , use the infrared modality target response R 1t to predict the target position. Subsequently, at the current target prediction position, use the scale-related filter to estimate the target size based on the HOG feature. When θ 1t <θ 1t-1 , predict the target state through the adaptive decision fusion of the visible light modality and the infrared modality. The specific method is the same as the RGB-T multi-modal fusion target tracking mode.
[0065] When θ 1t ≥θ 1t-1When, in each frame of the infrared image, the displacement-related filter is trained according to spatio-temporal regularization and sparse target response constraints, and the scale-related filter is trained by an incrementally updated appearance model. The specific method is the same as that of the infrared single-modal target tracking mode. When θ 1t < θ 1t-1 When, the two-modal displacement-related filter is jointly trained according to spatio-temporal regularization and the consistency constraint of visible light and infrared sparse target responses, and the scale-related filter is trained by an incrementally updated appearance model in the visible light image. The specific method is the same as that of the RGB-T multi-modal fusion target tracking mode.
[0066] This application has multiple observation models and can flexibly switch between single-modal target tracking and RGB-T target tracking; spatio-temporal regularization is used to suppress overfitting and drastic changes of the correlation filter coefficients, and sparse target response constraints are used to weaken the multi-peak interference of the target response, effectively improving the robustness of the infrared single-modal target tracking observation model. The sparse target response consistency constraint is used to jointly train the visible light correlation filter and the infrared correlation filter to fully exploit the two-modal complementarity and effectively improve the infrared target tracking performance in challenging scenarios such as thermal crossover. The reliability of the infrared modality is judged based on the target response, and the visible light modality is adaptively introduced and decision fusion is achieved according to the modality reliability.
[0067] S102. When the aircraft conducts air-to-ground infrared time-sensitive target tracking, determine the corresponding observation model according to the tracking scenario; the observation models include: infrared single-modal observation model, RGB-T multi-modal fusion observation model, and visible light-assisted infrared observation model;
[0068] S103. Use the observation model corresponding to the tracking scenario to perform target tracking on the current video frame.
[0069] To verify the feasibility and effectiveness of this application, time-sensitive target tracking simulation tests were carried out on multiple groups of air-to-ground RGB-T video sequences, and the corresponding tracking results were obtained. The hardware platform of this application is a computer configured with an Intel(R) Core(TM) i5-8300 CPU @ 2.30H, and the software platform is MATLAB R2018a.
[0070] Figure 5 The tracking results of this application for 3 groups of air-to-ground RGB-T video sequences are given, where parts (a), (b), and (c) of 5 respectively adopt the infrared single-modal target tracking mode, the RGB-T multi-modal fusion target tracking mode, and the visible light-assisted infrared target tracking mode. From the experimental results, it can be seen that the method proposed in this application can stably and reliably track the air-to-ground infrared time-sensitive targets in the sequence.
[0071] In this application, all actions of obtaining signals, information or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0072] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0073] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. An air-to-ground infrared time-sensitive target tracking method based on multi-mode correlation filtering, characterized in that, The method for tracking airborne infrared time-sensitive targets based on multi-modal correlation filtering includes: Constructing an infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model according to the tracking scenario; the infrared single-modal observation model calculates the target response map using a displacement correlation filter, estimates the target size using a scale correlation filter; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model; the RGB-T multi-modal fusion observation model calculates the target responses of the visible light and infrared modalities respectively using a displacement correlation filter; calculates the tracking reliability of each modality using the target responses; adaptively decides on the fusion to predict the target state according to the modality reliability; and jointly trains the displacement correlation filters of the visible light and infrared modalities according to spatio-temporal regularization and the consistency constraints of the visible light and infrared sparse target responses, and trains the scale correlation filter through an incrementally updated appearance model in the visible light image; the visible light-assisted infrared observation model calculates the target response of the infrared modality using a displacement correlation filter; determines the tracking reliability using the target response of the infrared modality; predicts the target state according to the tracking reliability; and trains the displacement correlation filter according to spatio-temporal regularization and sparse target response constraints, and trains the scale correlation filter through an incrementally updated appearance model. When the aircraft performs airborne infrared time-sensitive target tracking, determine the corresponding observation model according to the tracking scenario; the observation models include: an infrared single-modal observation model, an RGB-T multi-modal fusion observation model, and a visible light-assisted infrared observation model. Use the observation model corresponding to the tracking scenario to perform target tracking on the current video frame.
2. The method for tracking airborne and ground infrared time-sensitive targets based on multi-mode correlation filtering according to claim 1, wherein The infrared single-modal observation model calculates the target response map using a displacement correlation filter and estimates the target size using a scale correlation filter, specifically including: Extract the histogram of oriented gradients features, gray-scale features, and motion features of the target search area in the infrared image. Estimate the target response map using the displacement correlation filter based on the histogram of oriented gradients features, gray-scale features, and motion features. Estimate the target size using the scale correlation filter based on the histogram of oriented gradients features.
3. The method for tracking air-ground infrared time-sensitive targets based on multi-mode correlation filtering according to claim 1, wherein The infrared single-modal observation model uses to train the displacement-related filter h 1t ; Among them, is the displacement correlation filter of the d-th dimension of the t-th frame, is the d-th dimensional image block vectorized feature map of the t-th frame for training the displacement correlation filter h 1t , D1 is the displacement correlation filtering dimension, is the displacement correlation filter of the d-th dimension of the (t - 1)-th frame, is the vectorized ideal target response, is the set of real numbers, * is the correlation operation operator, ⊙ is the Hadamard product, is the spatial regularization constraint weight, λ1 is the coefficient of the spatial regularization constraint term, λ2 is the coefficient of the temporal regularization constraint term, λ3 is the coefficient of the sparse target response constraint term, T is the vector length of the vectorized feature map of each channel, is the F-norm Frobenius, || ||1 is the vector l1 norm.
4. The method for tracking airborne and ground infrared time-sensitive targets based on multi-mode correlation filtering according to claim 1, characterized in that, The determination of the tracking reliability using the target response of the infrared modality specifically includes: Using the formula θ 1t = APCE(R 1t ) max(R 1t ) to determine the tracking reliability θ 1t ; Among them, APCE(R 1t ) is the average peak correlation energy ratio of the infrared mode target response R 1t , and max(R 1t ) is the maximum value of the infrared mode target response R 1t .
5. The method for tracking air-ground infrared time-sensitive targets based on multi-mode correlation filtering according to claim 4, wherein The prediction of the target state according to the tracking reliability specifically includes: When θ 1t ≥ θ 1t-1 , the infrared modal target response R 1t is used to predict the target position. At the currently predicted target position, a scale correlation filter is used to estimate the target size based on the histogram of oriented gradient features; where θ 1t-1 is the tracking reliability of the (t - 1)-th frame. When θ 1t <θ 1t-1 , the RGB-T multi-modal fusion observation model is used to predict the target state.