An Asymmetric Hierarchical Feature Fusion Method for RGBT Object Tracking

By adopting an asymmetric hierarchical feature fusion method in RGBT target tracking, using dual-stream network and specific modal feature extraction module, the problem of poor tracking effect in harsh scenarios in the prior art is solved, and higher tracking accuracy and robustness are achieved.

CN116503442BActive Publication Date: 2025-06-03SICHUAN CHAOYIHONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310387925.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-12
Publication Date
2025-06-03
Estimated Expiration
2043-04-12

AI Technical Summary

Technical Problem

The existing RGBT target tracking methods have poor tracking effects in harsh scenarios such as extreme light and smoke, mainly because a single network structure cannot fully utilize the heterogeneity of visible light and thermal infrared modes.

Method used

Asymmetric hierarchical feature fusion method is adopted, and the features of visible light and thermal infrared images are extracted respectively by extending the backbone network of VGG-M to a dual-stream network. In the asymmetric hierarchical fusion structure, the visible light branches are mainly used and thermal infrared branches are supplemented. The Ghost modal feature extraction module and SKNet feature adaptive aggregation and enhancement are used for feature aggregation and enhancement.

Benefits of technology

By fully leveraging the heterogeneity of visible light and thermal infrared modes, the accuracy and robustness of target tracking are improved, especially in extreme conditions, where the target can be positioned and tracked more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503442B_ABST
    Figure CN116503442B_ABST
Patent Text Reader

Abstract

The present invention discloses an RGBT object tracking method with asymmetric hierarchical feature fusion, belonging to the technical fields of computer vision and object tracking. The method adopts different branch designs for visible light and thermal infrared modalities by introducing an asymmetric hierarchical feature fusion structure, with the visible light as the main branch and the thermal infrared as the auxiliary branch, making full use of the heterogeneity of the visible light and thermal infrared modalities. At the same time, when enhancing the visible light branch, the deep and shallow features of the thermal infrared modality are fully utilized to give full play to the complementarity of the two modalities. Moreover, in order to utilize more shallow features, the method of the present application adopts the way of hierarchical feature fusion, thereby making the recognition result more accurate. After obtaining the tracking result by regressing the bounding box through the candidate box regressor, the Alpha-Refine module is used for further refinement processing to make the tracking result more accurate and improve the performance of the tracker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an RGBT object tracking method with asymmetric hierarchical feature fusion, belonging to the technical fields of computer vision and object tracking. Background Art

[0002] Visual object tracking is one of the important tasks in computer vision, and object tracking has important applications in fields such as autonomous driving and video surveillance. The current single-object tracking task has poor tracking effects for extreme lighting, smoke and other harsh scenarios, because only the visible light modality is used in the single-object tracking task. With the development of thermal infrared technology and the advantages of infrared imaging being insensitive to light and able to penetrate smoke, etc., it can, to a certain extent, make up for the deficiencies of the visible light modality in tracking under harsh conditions. Therefore, RGBT (RGB-Thermal) object tracking based on the fusion of visible light and infrared images has attracted wide attention.

[0003] The key issue of RGBT object tracking is to make full use of the visible light and thermal infrared modalities and the fusion of the two modalities. Existing RGBT object tracking methods are mainly divided into discriminative fusion and combinative fusion. Discriminative fusion is to fuse by mining discriminative information of different modalities, and combinative fusion is to combine information of different modalities through different fusion methods. Most of the existing algorithms use the same network structure to process the visible light and thermal infrared modalities. However, due to different imaging principles of visible light and thermal infrared images, the visible light image contains more color texture information, and the thermal infrared image does not contain color texture information, but contains the appearance contour information of the object and the deep semantic information to help object classification. The thermal infrared image has good complementarity to the visible light image under extreme conditions such as lack of light. However, most of the existing algorithms using the same network structure for visible light and thermal infrared images cannot make full use of the heterogeneity of the two modalities, so the object tracking accuracy needs to be further improved. Summary of the Invention

[0004] To solve the problem that the visible light and thermal infrared modalities are heterogeneous, resulting in low target tracking accuracy, the present invention provides an RGBT target tracking method with asymmetric hierarchical feature fusion. Different branches can be designed according to the different characteristics of the two modalities, improving the performance of RGBT target tracking and meeting the design requirements of actual engineering systems. Based on the MDNet framework, the backbone network of VGG-M is first extended to a two-stream network to extract the features of visible light and thermal infrared images respectively. In the asymmetric hierarchical fusion structure, the visible light branch is the main branch and the thermal infrared branch is the auxiliary branch. For each layer, a Ghost-based modal feature extraction module is used to extract features, and then an SKNet-based feature adaptive aggregation module is used to aggregate the features of the two modalities. The obtained aggregated features are then enhanced through an enhanced fusion module for each layer of the visible light main branch, thereby realizing hierarchical feature fusion during the process of feature enhancement, making full use of the deep and shallow features of the thermal infrared modality. Compared with the existing method of fusing after enhancement, the method of the present application can utilize more shallow features, thus further improving the tracking accuracy.

[0005] An RGBT target tracking method with asymmetric hierarchical feature fusion, the method is implemented based on an RGBT target tracking network with asymmetric hierarchical feature fusion. The backbone network of the RGBT target tracking network with asymmetric hierarchical feature fusion adopts the first three layers conv1-conv3 of VGG-M and extends it to a two-stream network, with the visible light branch as the main branch and the thermal infrared branch as the auxiliary branch, and three fully connected layers are connected subsequently; the method includes:

[0006] Step S1, input the visible light image and the thermal infrared image corresponding to the current video frame into the RGBT target tracking network with asymmetric hierarchical feature fusion;

[0007] Step S2, respectively obtain the features of the visible light modality of the i-th layer and the features of the thermal infrared modality as well as the auxiliary visible light modality features of the i-th layer and the auxiliary thermal infrared modality features

[0008] Step S3, obtain the aggregated features of the i-th layer according to the auxiliary visible light modality features and the auxiliary thermal infrared modality features

[0009] Step S4, for each layer, interactively fuse the features of the visible light modality and the aggregated features through a cross-attention mechanism to obtain enhanced fusion features as the input of the next layer;

[0010] Step S5, for the third layer, aggregate the features with the features of the visible light modality and the features of the thermal infrared modality to perform separate interactive fusion respectively, and obtain enhanced fusion features respectively;

[0011] Step S6, input the two enhanced fusion features obtained in the third layer into three fully connected layers, discriminate the foreground and background of the current video frame, locate the target position, and achieve target tracking.

[0012] Optionally, in the two-stream network, for each layer, it includes a backbone structure, an asymmetric hierarchical fusion structure, and a feature enhancement fusion module FEF; the backbone structure is used to obtain the features of the visible light modality and the features of the thermal infrared modality The asymmetric hierarchical fusion structure is used to obtain aggregated features The feature enhancement fusion module FEF splits the encoder and decoder in the traditional Transformer structure to form two encoders and one decoder structure, and takes the features of the visible light modality obtained by the backbone network and the aggregated features obtained by the asymmetric hierarchical fusion structure as the inputs of the two encoders respectively for enhancement, and then uses the decoder structure to interactively fuse the features enhanced by the encoder through the cross-attention mechanism to obtain the enhanced fusion features of the i+1 layer

[0013] For the third layer, an additional feature enhancement fusion module FEF is set up to enhance the features of the thermal infrared modality and perform interactive fusion with the aggregated features to obtain enhanced fusion features

[0014] Input the enhanced fusion features obtained in the third layer and into three fully connected layers to discriminate the foreground and background of the current video frame and locate the target position.

[0015] Optionally, the asymmetric hierarchical fusion structure includes a Ghost-based modal feature extraction module MFE and a SKNet-based feature adaptive aggregation module FAG; the Ghost-based modal feature extraction module MFE is used to obtain the auxiliary visible light modal features of the i-th layer and the auxiliary thermal infrared modal features The feature adaptive aggregation module based on SKNet is used to obtain the i-th layer of aggregated features according to the auxiliary visible light modality features and the auxiliary thermal infrared modality features to obtain the i-th layer of aggregated features

[0016] Optionally, when obtaining the i-th layer of auxiliary visible light modality features and the auxiliary thermal infrared modality features in step S2, it is calculated according to the following formula:

[0017]

[0018] where is the feature map input to the modality feature extraction module MFE for each layer, is the new feature map obtained by first passing through a common convolution, is another feature map obtained through a linear operation, is to concatenate the two obtained feature maps to get the final feature map, that is and ε represents the concatenation operation, l represents the number of layers, represents the convolution operation The convolution kernel sizes of are 5*5, 3*3, and 1*1 respectively; represents the grouped convolution operation, the number of groups is C 2 / 4, and the convolution kernel size is 3*3.

[0019] Optionally, when obtaining the i-th layer of aggregated features in step S3, it is calculated according to the following formula:

[0020]

[0021]

[0022] fc = f 2 (γ(f 1 ))

[0023] where, w rgb , w t represent the weights of the adaptive aggregation module FAG, α represents the global average pooling operation, f 1 and f 2 represent the 1*1 convolution operation, γ represents the ReLu activation function, and φ represents the softmax function.

[0024] Optionally, step S4 includes:

[0025] The features obtained for each layer of the visible light backbone network Obtain aggregated features through an asymmetric hierarchical fusion structure Obtain the q, k, and v vectors through linear transformation and use them as the inputs of the two Encoders respectively;

[0026] Enhance through the self-attention mechanism, and the attention weight matrix Atten(q,k,v) = softmax(qk T / τ)v, where τ = 30, and then use the residual to add the vector v to the original vector;

[0027] In the Decoder part, the features obtained by the visible light modality through the Encoder are used as v, and the features obtained by the aggregated features of the thermal infrared modality through the Encoder are used as q; through the cross-attention mechanism, and then add the vector v to the original vector to obtain the enhanced and fused features

[0028] Optionally, the convolutional kernel sizes of the first three layers conv1-conv3 of the backbone network are 7*7, 5*5, and 3*3 respectively.

[0029] Optionally, the output dimensions of the three fully connected layers are 512, 512, and 2 respectively.

[0030] Optionally, the method further includes training the RGBT object tracking network with asymmetric hierarchical feature fusion using a publicly available dataset, and iteratively training and updating the network parameters through a loss function and a stochastic gradient descent algorithm during the training process.

[0031] Optionally, the publicly available dataset includes three RGBT tracking datasets: GTOT, RGBT234, and LasHeR.

[0032] The beneficial effects of the present invention are:

[0033] By introducing an asymmetric hierarchical feature fusion structure, different branch designs are adopted for the visible light and thermal infrared modalities, with the visible light as the main branch and the thermal infrared as the auxiliary branch, making full use of the heterogeneity of the visible light and thermal infrared modalities. At the same time, when enhancing the visible light branch, the deep and shallow features of the thermal infrared modality are fully utilized to give full play to the complementarity of the two modalities. Moreover, in order to utilize more shallow features, the method of hierarchical feature fusion is adopted in this application, thus making the recognition result more accurate. After obtaining the tracking result by regressing the bounding box through the candidate box regressor, the Alpha-Refine module is used for further refinement to make the tracking result more accurate and improve the performance of the tracker. Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0035] Figure 1 It is the implementation flowchart of the method of the present invention.

[0036] Figure 2 It is the architecture diagram of the RGBT target tracking network based on asymmetric hierarchical feature fusion provided by the present invention.

[0037] Figure 3 It is the overall network framework diagram of the present invention.

[0038] Figure 4 It is the schematic diagram of the module feature extraction module of the present invention.

[0039] Figure 5 It is the schematic diagram of the feature adaptive aggregation module of the present invention.

[0040] Figure 6 It is the schematic diagram of the feature enhancement fusion module of the present invention.

[0041] Figure 7A and 7B It is the comparison result diagram of the method of the present invention with APFNet, MANet, and SGT on the GTOT dataset. Figure 7A It is the accuracy diagram. Figure 7B It is the success diagram.

[0042] Figure 8A and 8B It is the comparison result diagram of the method of the present invention with APFNet, MANet, and SGT on the RGBT234 dataset. Figure 8A It is the accuracy diagram. Figure 8B It is the success diagram.

[0043] Figure 9A and 9B It is the comparison result diagram of the method of the present invention with APFNet, MANet, and SGT on the LasHeR dataset. Figure 9A It is the accuracy diagram. Figure 9B It is the success diagram. Detailed implementation manners

[0044] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0045] Introduction to the network structure involved in this application:

[0046] 1. MDNet Algorithm

[0047] MDNet is an algorithm proposed based on multi - domain learning. It regards each video sequence as a separate domain. The main architecture of MDNet is shown as follows Figure 3 MDNet's backbone network is based on the VGG - M architecture. Its conv1 - conv3, fc4, and fc5 layers are shared layers, and the last layer, fc6 layer, is domain - specific. MDNet models the domain - independent layers in the shared layers to obtain a general feature representation. For the domain - specific layer fc6, a classifier that distinguishes the target from the background is trained for the purpose of tracking.

[0048] MDNet crops the sampled samples to a size of 107 * 107, and then inputs them into the network shown as follows Figure 2 to output the confidence scores of the samples. In MDNet, a hard negative mining strategy is adopted to solve the offset problem. In each iteration, the 96 negative samples with the highest scores and 32 positive samples are taken to iteratively train and update the model parameters through the following loss function and the stochastic gradient descent (SGD) algorithm:

[0049]

[0050] where N represents the number of samples in the training phase, y i represents the true label of the i - th sample, and p i represents the predicted probability value of the i - th sample.

[0051] 2. Feature Enhancement and Fusion Module Based on Transformer Structure

[0052] Transformer contains encoder and decoder structures. In the method of the present invention, the encoder and decoder are split to form two encoders and one decoder structure. Two encoders are used to enhance the two input features respectively, and the decoder structure is used to interact and fuse the features enhanced by the encoders. In the method of the present invention, a single - head attention mechanism is adopted for simplicity.

[0053] The encoder and decoder are composed of self - attention mechanism layers and feed - forward neural network layers. The input of the self - attention mechanism is three vectors, and then these three vectors are multiplied by three matrices to obtain three new vectors, and then the attention scores are calculated:

[0054]

[0055] Among them, three new vectors calculated from Q, K, and V, d k represents the dimension of the vector.

[0056] The output of the self-attention layer is the input of the feed-forward neural network, and a matrix is input through the feed-forward neural network.

[0057] Both the encoder and decoder structures contain residual structures, that is, the input of each feed-forward network contains not only the output of the self-attention layer but also the original input.

[0058] Example 1:

[0059] This embodiment provides an RGBT target tracking method based on asymmetric hierarchical feature fusion, which is implemented based on an RGBT target tracking network with asymmetric hierarchical feature fusion. The backbone network of the RGBT target tracking network with asymmetric hierarchical feature fusion adopts the first three layers conv1-conv3 of VGG-M and expands it into a two-stream network, with the visible light branch as the main branch and the thermal infrared branch as the auxiliary branch, and three fully connected layers are connected subsequently; see Figure 1 , the method includes:

[0060] Step S1, input the visible light image and thermal infrared image corresponding to the current video frame into the RGBT target tracking network based on asymmetric hierarchical feature fusion;

[0061] Step S2, respectively obtain the features of the visible light modality of the i-th layer and the features of the thermal infrared modality as well as the auxiliary visible light modality features of the i-th layer and the auxiliary thermal infrared modality features

[0062] Step S3, obtain the aggregated feature of the i-th layer according to the auxiliary visible light modality features and the auxiliary thermal infrared modality features

[0063] Step S4, for each layer, interactively fuse the features of the visible light modality and the aggregated feature through the cross-attention mechanism to obtain an enhanced fusion feature as the input of the next layer;

[0064] Step S5, for the third layer, the aggregated feature is combined with the features of the visible light modality and the features of the thermal infrared modality ​Perform separate interactive fusion to obtain enhanced fusion features respectively;

[0065] Step S6: Input the two enhanced fusion features obtained in the third layer into three fully connected layers to discriminate the foreground and background of the current video frame, locate the target position, and achieve target tracking.

[0066] Embodiment 2:

[0067] This embodiment provides an RGBT target tracking method with asymmetric hierarchical feature fusion, including:

[0068] Step 1. Construct an RGBT target tracking network with asymmetric hierarchical feature fusion;

[0069] As Figure 2 shown, the constructed RGBT target tracking network with asymmetric hierarchical feature fusion is based on MDNet. The backbone network adopts the first three layers conv1-conv3 of VGG-M and expands it into a two-stream network, with the visible light branch as the main branch and the thermal infrared branch as the auxiliary branch.

[0070] In the two-stream network, for each layer, it includes a backbone structure, an asymmetric hierarchical fusion structure, and a feature enhancement fusion module FEF. Each layer of the backbone structure is used to obtain the features of the visible light modality and the features of the thermal infrared modality of each layer and the features of the thermal infrared modality The asymmetric hierarchical fusion structure includes a Ghost-based modal feature extraction module MFE and an SKNet-based feature adaptive aggregation module FAG. The Ghost-based modal feature extraction module MFE is used to obtain the auxiliary visible light modal features of the i-th layer and the auxiliary thermal infrared modal features The SKNet-based feature adaptive aggregation module is used to obtain the i-th layer aggregation feature according to the auxiliary visible light modal features and the auxiliary thermal infrared modal features The feature enhancement fusion module FEF splits the encoder and decoder in the traditional Transformer structure to form two encoders and one decoder structure, and uses the visible light modal features obtained by the backbone network and the aggregation features obtained by the asymmetric hierarchical fusion structure and the aggregation features obtained by the asymmetric hierarchical fusion structure as the inputs of the two encoders for enhancement respectively, and then uses the decoder structure to interactively fuse the features enhanced by the encoder through the cross-attention mechanism to obtain the enhanced fusion features of the (i + 1)-th layer and input them to the next layer.

[0071] For the last layer, in order to utilize the deep features of the thermal infrared image, a feature enhancement and fusion module is also added to the last layer of feature extraction. By inputting the features obtained from conv3 into an encoder for enhancement, and then inputting the aggregated and enhanced features and the enhanced thermal modality features obtained from the third layer into a decoder for interactive fusion to obtain the final enhanced fusion features. Finally, the two enhanced fusion features obtained from the third layer are input into three fully connected layers to discriminate the foreground and background of the current video frame, locate the target position, and achieve target tracking.

[0072] Specifically, the method includes:

[0073] (1.1) Input visible light and thermal infrared images, and then use Gaussian sampling to obtain positive and negative samples.

[0074] In this embodiment, 32 positive samples and 96 negative samples are obtained by sampling.

[0075] (1.2) The sampled samples are input into the backbone network for feature extraction to obtain the features of the visible light modality of the i-th layer and the features of the thermal infrared modality

[0076] Here, the asymmetric hierarchical feature fusion network is based on MDNet. The backbone network adopts the first three layers of VGG-M and expands it into a two-stream network. The convolutional kernel sizes of the three layers are 7*7, 5*5, and 3*3 respectively.

[0077] (1.3) Based on the Ghost-based modal feature extraction module MFE, obtain the auxiliary visible light modality features of the i-th layer and the auxiliary thermal infrared modality features Refer to Figure 4 , and the calculation is as follows:

[0078]

[0079] Where is the feature map input into the modal feature extraction module for each layer, is to first pass through a common convolution to obtain a new feature map, is to obtain another feature map through linear operation, is to concatenate the two obtained feature maps to obtain the final feature map, which is and ε represents the concatenation operation, l represents the number of layers, represents the convolution operation The convolution kernel sizes of Denotes a grouped convolution operation with the number of groups being C 2 / 4, and the convolution kernel size is 3*3.

[0080] (1.4) Obtain the aggregated feature of the i-th layer based on the Feature Adaptive Aggregation Module FAG of SKNet Refer to Figure 5 , and it is calculated as follows:

[0081]

[0082]

[0083] fc = f 2 (γ(f 1 ))

[0084] Among them, w rgb , w t represents the weight of the adaptive feature extraction module, α represents the global average pooling operation, f 1 and f 2 represent the 1*1 convolution operation, γ represents the ReLu activation function, and φ represents the softmax function.

[0085] (1.5) The Feature Enhancement and Fusion Module FEF obtains the enhanced fusion feature of the i+1-th layer Refer to Figure 6 , and proceed as follows:

[0086] (a) The features obtained by each layer of the visible light backbone network and the aggregated features obtained by the asymmetric hierarchical fusion structure Obtain the q, k, and v vectors through linear transformation and use them as the inputs of the two Encoders respectively.

[0087] (b) Enhance through the self-attention mechanism, and the attention weight matrix Atten(q,k,v) = softmax(qk T / τ)v, where τ = 30, and then use the residual to add the vector v to the original vector.

[0088] (c) In the Decoder part, the features obtained by the visible light modality through the Encoder are used as v, and the features obtained by the aggregated features through the Encoder are used as q. Through the cross-attention mechanism, and then add the vector v to the original vector to obtain the enhanced fusion feature

[0089] Step 2. Use the publicly available datasets RGBT234 and GTOT datasets to train the above network model.

[0090] (2.1) Load the pre-trained model parameters of MDNet on ImageNetvid to initialize the backbone network, FC4, and FC6 layers, and randomly initialize other convolutional and fully connected layers.

[0091] (2.2) Set the learning rate of the fully connected layer to 0.001, the learning rate of the fully connected layer to 0.0001, the weight decay to 0.0005, the momentum to 0.9, and train for 200 epochs.

[0092] Step 3. Input the video sequence to be tracked, load the trained model for online tracking, obtain the tracking results through the candidate regressor, and use the Alpha-Refine module for refinement to obtain the target state of the final tracking target.

[0093] (3.1) Load the trained model, replace the multi-domain branch in the fc6 layer with a single branch, and randomly initialize the parameter W of the fc6 layer fc6 .

[0094] (3.2) Use the first frame of the video to train the bounding box regressor and initialize the Alpha-Refine module.

[0095] (3.3) Gaussian sample 500 positive samples on the first frame and 5000 negative samples Update the parameter W of fc4-fc6 using the sampled positive and negative samples fc4-fc6 .

[0096] (3.4) Generate long-term update and short-term update samples U l , U s .

[0097] (3.5) Gaussian sample 256 candidate regions of the current frame Z t-1 centered on the previous frame Z t Obtain the candidate region scores through the above-trained network Select the top five with the highest scores and take their average as the tracking result of the current frame

[0098] (3.6) If Perform bounding box regression on the tracking result, then input the result of the bounding box regression into the Alpha-Refine module, use the first frame of the video sequence as the reference frame, extract features of the two input images through the backbone network, then fuse the extracted features in a pixel-wise manner, and then obtain a more accurate tracking result.

[0099] If Perform short-term updates using the short-term update samples U s to update W fc4-fc6 .

[0100] Perform long-term updates every 10 frames using the long-term update samples U l to update W fc4-fc6 .

[0101] Perform long-term and short-term updates of the samples U after each successful tracking l , U s , and repeat (3.5) until all frames are tracked.

[0102] To verify the effectiveness of the method of this application, it is further illustrated by experiments as follows:

[0103] 1. Simulation conditions

[0104] Here, three RGBT tracking datasets, namely GTOT, RGBT234, and LasHeR, are adopted.

[0105] The GTOT dataset contains 50 visible light and thermal infrared video sequences, including 7 attribute categories. The specific attribute categories are occlusion (OCC), large-scale deformation (LSV), fast motion (FM), low illumination (LI), thermal crossover (TC), small target (SO), and deformation (DEF).

[0106] The RGBT234 dataset contains 234 visible light and thermal infrared video sequences, including 12 attribute categories. The specific attribute categories are no occlusion (NO), partial occlusion (PO), heavy occlusion (HO), low illumination (LI), low resolution (LR), thermal crossover (TC), deformation (DEF), fast motion (FM), scale variation (SV), motion blur (MB), camera movement (CM), and background clutter (BC), fully considering the challenges in various environments.

[0107] The LasHeR dataset contains 1224 aligned RGBT video sequences, including 19 attribute categories. The specific attribute categories are no occlusion (NO), partial occlusion (PO), total occlusion (TO), transparent occlusion (HO), out of view (OV), low illumination (LI), low resolution (LR), deformation (DEF), background clutter (BC), thermal crossover (TC), motion blur (MB), camera movement (CM), frame loss (FL), fast motion (FM), scale variation (SV), abrupt illumination variation (AIV), and aspect ratio change (ARC). Here, 245 video sequences are selected for experiments.

[0108] To verify the feasibility and effectiveness of the method proposed in the present invention, an asymmetric hierarchical feature fusion RGBT object tracking method (AHFNet) of the method of the present application is compared with three other existing methods, which are as follows:

[0109] Attribute-based progressive fusion network for RGBT tracking (APFNet), which can be referred to the introduction in "Attribute-Based Progressive Fusion Network for RGBT Tracking". Based on five specific attribute categories, including thermal crossing (TC), illumination variation (IV), scale variation (SV), occlusion (OCC), and fast motion (FM), five branches are designed to extract corresponding features for tracking.

[0110] Multi-adapter RGBT object tracking method (MANet), which can be referred to the introduction in "Multi-Adapter RGBT Tracking". A method based on adapters is designed, and three different adapters are designed to represent the target object well with general features, modality-specific features, and instance-aware features respectively to cope with different tracking challenges.

[0111] Graph learning-based RGBT object tracking method (SGT), which can be referred to the introduction in "Weighted Sparse Representation Regularized Graph Learning for RGB-T Object Tracking". A graph-based RGBT tracking method is proposed, and the modality weights, sparse representation, and graph are jointly learned to achieve a robust tracking effect.

[0112] The experimental running environment is as follows: Ubuntu 16.04, CUDA 10.0, python = 3.7, and the GPU is NVIDIA RTX 2080Ti. The present invention uses the precision rate (PR) and success rate (SR) in one-pass evaluation (OPE) as the evaluation indexes for quantitative analysis. Among them, PR represents the percentage of all frames in which the distance between the center point of the tracking result and the center point of the true position is less than the threshold. The larger the PR value, the more accurate the tracking. The threshold is set to 5 pixels in the GTOT dataset and 20 pixels in the RGBT234 and LasHeR datasets. SR is the percentage of images in which the intersection over union of the predicted bounding box and the true bounding box is greater than the threshold, and the score of SR is calculated by the area under the SR curve.

[0113] 2. Simulation results

[0114] The method of the present invention (AHFNet) is compared with the APFNet, MANet, and SGT methods.

[0115] Figure 7A And Figure 7B is a comparison graph of the results of AHFNet with APFNet, MANet, and SGT on the GTOT dataset, Figure 7A is the precision graph, Figure 7B is the success graph. From Figure 7A and Figure 7B it can be seen that the precision of the method of the present invention is 1.7% higher than that of the APFNet method, 2.8% higher than that of the MANet method, and 7.1% higher than that of the SGT method. \

[0116] Figure 8A And Figure 8B is a comparison graph of the results of AHFNet with APFNet, MANet, and SGT on the RGBT234 dataset, Figure 8A is the precision graph, Figure 8B is the success graph. From Figure 8A and Figure 8B it can be seen that the precision of the method of the present invention is comparable to that of the APFNet method, 5.2% higher than that of the MANet method, and the success rate is 3.2% higher than that of the APFNet method. \

[0117] Figure 9A And Figure 9B is a comparison graph of the results of AHFNet with APFNet, MANet, and SGT on the LasHeR dataset, Figure 9A is the precision graph, Figure 9B is the success graph. From Figure 9A and Figure 9B it can be seen that the precision of the method of the present invention is 2.7% higher than that of the APFNet method, 7.2% higher than that of the MANet method, and 20% higher than that of the SGT method. \

[0118] It can be clearly seen from the experimental result graphs that the tracking accuracy of the method of the present invention is better than that of the APFNet, MANet, and SGT methods. SGT is a traditional RGBT tracking method that extracts features manually, so its tracking results are relatively poor compared to the other three deep learning-based algorithms. MANet adopts a symmetric structure and does not fully utilize the heterogeneity of the visible light and infrared modalities, so its effect is relatively worse compared to the asymmetric structure adopted by the present method. In APFNet, five corresponding branches are designed for five category attributes. The GTOT dataset is relatively small and only contains 7 attribute categories, and the RGBT234 dataset contains 12 attribute categories. However, several categories are subdivisions of large attribute categories. For example, occlusion is divided into no occlusion (NO), partial occlusion (PO), and heavy occlusion (HO). Therefore, APFNet basically covers most of the attribute categories in these two datasets, so its tracking effect is relatively good. However, the present method can achieve a result slightly better than APFNet without designing attribute-based branches. However, on the LasHeR dataset, since the LasHeR dataset is a large dataset containing 19 attribute categories, APFNet cannot fully adapt to all scenarios. However, the present method has good robustness on the LasHeR dataset, indicating the effectiveness of the method of the present application.

[0119] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as a CD or a hard disk, etc.

[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An RGBT object tracking method based on asymmetric hierarchical feature fusion, characterized in that, the method is implemented based on an RGBT object tracking network with asymmetric hierarchical feature fusion. The backbone network of the RGBT object tracking network with asymmetric hierarchical feature fusion adopts the first three layers conv1 - conv3 of VGG - M and expands it into a two - stream network, with the visible - light branch as the main branch and the thermal - infrared branch as the auxiliary branch, and then connects three fully - connected layers; the method includes: Step S1, input the visible - light image and the thermal - infrared image corresponding to the current video frame into the RGBT object tracking network based on asymmetric hierarchical feature fusion; Step S2: Obtain the features of the visible light modality of the i-th layer, the features of the thermal infrared modality, the auxiliary visible light modality features of the i-th layer, and the auxiliary thermal infrared modality features respectively based on the visible light image and the thermal infrared image and the features of the thermal infrared modality as well as the auxiliary visible light modality features of the i-th layer and the auxiliary thermal infrared modality features Step S3, according to the auxiliary visible light modality features and the auxiliary thermal infrared modality features obtain the aggregated feature of the i-th layer Step S4. For each layer, interactively fuse the features in the visible light modality and the aggregated features through a cross-attention mechanism to obtain enhanced fused features as the input for the next layer; Step S5. For the third layer, aggregate the features with the features of the visible light modality and the features of the thermal infrared modality for separate interactive fusion to obtain enhanced fusion features respectively; Step S6, input the two enhanced fusion features obtained in the third layer into the three fully - connected layers, discriminate the foreground and background of the current video frame, locate the target position, and achieve object tracking; In the two-stream network, for each layer, it includes a backbone structure, an asymmetric hierarchical fusion structure, and a feature enhancement fusion module FEF; the backbone structure is used to obtain features of the visible light modality and features of the thermal infrared modality The asymmetric hierarchical fusion structure is used to obtain aggregated features The feature enhancement fusion module FEF splits the encoder and decoder in the traditional Transformer structure to form two encoders and one decoder structure, and takes the features of the visible light modality obtained by the backbone network and the aggregated features obtained by the asymmetric hierarchical fusion structure as the inputs of the two encoders for enhancement respectively, and then uses the decoder structure to interact and fuse the features enhanced by the encoders through the cross-attention mechanism to obtain the enhanced fusion features of the (i + 1)-th layer For the third layer, another feature enhancement and fusion module FEF is set up to enhance the features of the thermal infrared modality and then interact and fuse with the aggregated features to obtain enhanced fusion features after enhancement and interact and fuse with the aggregated features The enhanced fusion features obtained in the third layer and are input into three fully connected layers to discriminate the foreground and background of the current video frame and locate the target position; The asymmetric hierarchical fusion structure includes a Ghost-based modal feature extraction module MFE and an SKNet-based feature adaptive aggregation module FAG; the Ghost-based modal feature extraction module MFE is used to obtain the i-th layer of auxiliary visible light modal features and auxiliary thermal infrared modal features The SKNet-based feature adaptive aggregation module is used to obtain the i-th layer of aggregated features according to the auxiliary visible light modal features and auxiliary thermal infrared modal features ​ 2. The method according to claim 1, characterized in that, In step S2, the visible light modal features of the i-th layer of assistance and the thermal infrared modal features of assistance are calculated and obtained according to the following formula: Among them is the feature map input to the modal feature extraction module MFE for each layer, is the new feature map obtained by first passing through a common convolution, is another feature map obtained through a linear operation, is the final feature map obtained by concatenating the two obtained feature maps, that is and ε represents the concatenation operation, l represents the number of layers, represents the convolution operation The convolution kernels of represent the grouped convolution operation, and the number of groups is C 2 / 4, and the convolution kernel size is 3*3.

3. The method according to claim 2, characterized in that, When obtaining the i-th layer of aggregated features in step S3 it is calculated according to the following formula: fc = f 2 (γ(f 1 )) Among them, w rgb , w t represents the weight of the adaptive aggregation module FAG, α represents the global average pooling operation, f 1 and f 2 represent the 1*1 convolution operation, γ represents the ReLu activation function, and φ represents the softmax function.

4. The method according to claim 3, characterized in that, the step S4 includes: The features obtained from each layer of the visible light backbone network and the aggregated features obtained from the asymmetric hierarchical fusion structure Obtain q, k, and v vectors through linear transformation and use them as the inputs of the two Encoders respectively; Enhanced by the self-attention mechanism, the attention weight matrix Atten(q, k, v) = soft max(qk T / τ)v, where τ = 30, and then the vector v is added to the original vector using the residual; In the Decoder part, the features obtained by the visible light modality through the Encoder are used as v, and the features obtained by the aggregated features of v and k through the Encoder are used as q; through the cross-attention mechanism, then the vector v is added to the original vector to obtain the enhanced fusion features 5. The method according to claim 1, characterized in that, the convolution kernel sizes of the first three layers conv1 - conv3 of the backbone network are 7*7, 5*5, and 3*3 respectively.

6. The method according to claim 1, characterized in that, the output dimensions of the three fully - connected layers are 512, 512, and 2 respectively.

7. The method according to claim 1, characterized in that, the method further includes training the RGBT object tracking network with asymmetric hierarchical feature fusion using a public dataset, and iteratively training and updating the network parameters through a loss function and a stochastic gradient descent algorithm during the training process.

8. The method according to claim 7, characterized in that, the public dataset includes three RGBT tracking datasets, namely GTOT, RGBT234, and LasHeR.