RGBT target tracking method combining asymmetric enhancement and interactive fusion
By combining asymmetric enhancement and interactive fusion methods, the problem of insufficient modal fusion accuracy in complex environments of existing RGBT trackers is solved, and higher image quality and tracking accuracy are achieved to meet the needs of complex scenarios.
Patent Information
- Application Number
- CN202411910627.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-16
AI Technical Summary
Existing RGBT trackers cannot effectively balance performance and performance in complex environments, and the modal fusion accuracy is insufficient, making it difficult to meet the tracking target requirements in complex environments.
Using a method combining asymmetric enhancement and interactive fusion, the features of visible light and thermal infrared modes are extracted through the backbone network, image enhancement is performed separately, and information interaction and fusion between modes is performed through the multi-head cross-attention mechanism, and finally classification and regression is used using the prediction head.
Improve image quality, enhance the accuracy of modal fusion, improve tracking accuracy and robustness, and better adapt to complex tracking scenarios.
Smart Images

Figure CN120013990A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an RGBT target tracking method combining asymmetric enhancement and interactive fusion. Background Art
[0002] In complex environments such as low lighting, haze, and darkness, visible light images often fail to provide enough details, while thermal infrared images can highlight targets in dark or low-light conditions, but usually lack color information and details. To solve this problem, RGBT (visible light and thermal infrared) trackers came into being. By combining visible light and thermal infrared images, more robust target tracking can be achieved in all-weather and all-day conditions. This fusion can make up for the shortcomings of each modality, making target tracking more accurate in complex environments.
[0003] However, existing RGBT trackers often have certain limitations in modal fusion. Many trackers fail to effectively fuse the information of the two modalities, or ignore the interaction between the modalities during the fusion process. For example, the processing of modal fusion usually relies on simple weighted average or early fusion strategies, which may fail to fully exploit the complementary advantages of visible light and thermal infrared modalities, resulting in insufficient tracking accuracy and robustness. Therefore, the existing technology has the problem of inaccurate target tracking in complex environments.
[0004] Enlightengan proposed a global-local discriminant structure, self-regularized perceptual loss fusion and attention mechanism for image enhancement. However, this method failed to fully consider the comprehensive enhancement of channel dimension and spatial dimension. Optimizing a single dimension alone may lead to distortion of detail or structural information in some cases, affecting image quality. STANet introduced a guided filter to decompose the image into structural part and texture part to enhance global and local features. PSTLFusion extracts coarse-grained and fine-grained information through texture learning and models the internal structure of the image in detail. However, over-reliance on fine-grained information may lose the global structure of the image or cause inconsistency between details and the overall structure. EDFNet designed detail information guidance (DiG) and structural information enhancement (SiE). DiG optimizes the details and texture features of RGB features through axial details. SiE uses cumulative addition of features to enhance structural features. However, the common problems of these methods are that they may amplify noise or lead to over-enhancement, and often ignore the comprehensive enhancement of channel dimension and spatial dimension, resulting in insignificant local detail enhancement or incomplete global structure.
[0005] Venkatachalam proposed to solve the problems of low contrast and blurred details by attenuated color channel correction and detail-preserving contrast enhancement, but this solution requires a relatively complex parameter adjustment and calculation process, especially in the case of extremely low illumination or large noise, it is difficult to ensure the stability of the enhancement effect. EEMEFN uses a two-stage method of edge-enhanced multi-exposure fusion network to enhance extremely low-illumination images. This method requires multi-stage processing and involves multiple exposures and edge enhancement, so the computational complexity is high and it is not suitable for real-time applications. EC-CNN proposed an edge-conditional convolutional neural network and designed a feature transformation layer to adaptively enhance edge information. The feature transformation layer increases the amount of calculation, especially in complex scenes where edge information needs to be enhanced, which may lead to a decrease in computational efficiency. Srinivasu proposed a sharpening algorithm technology to overcome the low contrast problem of the fused image. R2rnet proposed to use the spatial information of the image to improve the contrast and also use the frequency information to retain details. Although the above methods have achieved certain results, they still cannot maintain high image quality for complex scenes or when the background changes greatly.
[0006] In order to integrate the different features of the two modalities, RITFusion consists of an intra-modal self-attention (IMSA) block, a modal reinforcement (MR) block, and an inter-modal interactive attention (IMIA) block to achieve information association and interaction between modalities, but it fails to fully consider the different needs of dynamic scene changes for modal information, and may not be able to adjust automatically in some scenes. AITFuse alternately combines local and global relationships through collaborative learning of CNN and Transformer, proposes a cascaded token and channel Transformer architecture with different attention mechanisms to model remote contexts, and communicates features across different tokens and channels in an interactive manner. However, the local details may not be processed finely enough, especially when the image target is small or the local changes are large, which may affect the tracking accuracy. APFNet proposes an attribute-based progressive fusion network, adopts a specific fusion strategy, and integrates three independent encoders (based on self-attention mechanism) and decoders (based on cross-attention mechanism) in each fusion branch. Although the above trackers rely on complex designs to achieve good performance in tracking accuracy, they still have poor robustness and poor adaptability to be optimized and improved.
[0007] In summary, existing technologies cannot effectively cope with dynamic changes in complex environments while balancing performance and expression. The modal fusion accuracy is insufficient and it is difficult to meet the actual needs of tracking targets in complex environments. Summary of the invention
[0008] The present invention provides an RGBT target tracking method combining asymmetric enhancement and interactive fusion to solve the problems in the prior art that the tracker cannot effectively cope with dynamic changes in complex environments while balancing performance and expression, and the modal fusion accuracy is insufficient.
[0009] In order to achieve the above object, the technical solution provided by the present invention is as follows: a RGBT target tracking method combining asymmetric enhancement and interactive fusion, comprising the following four steps:
[0010] Step 1: Feature extraction: Use the backbone network to extract feature information from the visible light modality and thermal infrared modality respectively;
[0011] Step 2, feature enhancement: The feature information enters the feature enhancement step, in which the feature information is sent to a visible light image enhancement module and a thermal infrared image enhancement module in parallel for processing, and two modes are output: an enhanced visible light image and a thermal infrared image;
[0012] Step 3: Interactive fusion: Add the two modalities after self-enhancement, and use the added modality as the key K' and value V' of multi-head cross attention, and the other modality as the query Q' for fusion;
[0013] Step 4: Prediction: Use the prediction head for classification and regression.
[0014] Furthermore, the visible light image enhancement module in the above step 2 includes two parts: detail-semantic enhancement in the channel dimension and detail-semantic enhancement in the spatial dimension. The detail-semantic enhancement in the spatial dimension captures information of different scales through multi-scale feature extraction, and then models the image context relationship through the spatial self-attention mechanism. While modeling the spatial self-attention mechanism, a spatial weighting mechanism is used to design spatial weighting to enhance the spatial information of the image.
[0015] Furthermore, the above multi-scale feature extraction includes 1 1×1 convolution, 3 3×3 dilated convolutions with different expansion rates and global average pooling. The feature maps extracted from different scales are connected in the channel dimension to obtain the enhanced detail feature map F, which is calculated as follows:
[0016] F=Cat(f3(X),f4(X),f5(X),f6(X),U(GAP(X)) (4)
[0017] Among them, Cat represents the connection operation, f3 represents 1×1 convolution, f4, f5, f6 represent dilation rates of 6, 12, and 18 respectively, U represents the upsampling operation, and F represents the enhanced detail features.
[0018] Furthermore, the specific calculation formula based on the spatial self-attention mechanism modeling is:
[0019] Q=W q F (5)
[0020] K=W k F (6)
[0021] R=Q T K (7)
[0022] R'=S(Maxpool(Q T K) (8)
[0023] Through the feature F and the corresponding weight matrix W v Perform a linear transformation to obtain V∈R C×H×W :
[0024] V=W v F (9)
[0025] Take the dot product of the feature map V and R' to get the output V of the spatial self-attention modeling module r The mathematical expression is:
[0026] V r =R'⊙V (10)
[0027] Furthermore, the specific calculation formula of the spatial weighting in the spatial weighting mechanism design is:
[0028] α=LRN(ReLU(Conv)) (11)
[0029] F1=α(max(F)) (12)
[0030] w=S(Conv(F1)) (13)
[0031] The mathematical expression of the final enhanced feature V is:
[0032]
[0033] Furthermore, the thermal infrared image enhancement module in the above step 2 includes an edge information enhancement module, an adaptive contrast enhancement module and a texture information enhancement module. The edge information enhancement module is designed based on the Laplace operator and the Sobel operator. The thermal infrared features are first sent to the edge information enhancement module, and then pass through the adaptive contrast enhancement module. The enhanced information is then passed through the texture information enhancement module for texture enhancement in the horizontal and vertical directions respectively.
[0034] Furthermore, the edge information enhancement module adopts a dual-branch structure, and the upper branch is used for thermal infrared features T hThe Laplacian operator is used to highlight the edge features of the image by calculating the second-order derivative of the image. After that, a convolution block (Conv-BN-ReLU) is passed to finally add the edge feature maps obtained from the two branches element by element to obtain the final edge feature map T e :
[0035]
[0036] Where β=ReLU(BN(Conv)), La represents the Laplace operator, So represents the Sobel operator, and BN represents local response normalization.
[0037] Furthermore, the specific process of the above-mentioned adaptive contrast enhancement module is as follows:
[0038] First, input the edge-enhanced thermal infrared image feature T e , calculate the standard deviation of each position (i, j) in the feature, the calculation formula is:
[0039]
[0040] Where w and h represent the width and height of the thermal infrared image, respectively, and μ ij , σ ij Respectively represent the mean and standard deviation of the image, P(i,j) represents the pixel value of the i-th row and j-th column;
[0041] Then there are two 3×3 convolutional layers, ReLU and Sigmoid activation functions to obtain the contrast weights of the feature map;
[0042] Finally, the edge feature T e Multiplying with the contrast weights yields the enhanced feature T s :
[0043] T s =T e ×S(A(σ ij (T e )) (17)
[0044] Where A = Conv(ReLU(Conv)).
[0045] Furthermore, the specific steps of the above texture enhancement module are:
[0046] Feature T s First, texture enhancement is performed in the horizontal direction of the thermal infrared image. Atrous convolution, element-by-element convolution and ordinary convolution are used to connect the images obtained after different convolutions, and then dimensionality reduction is performed to obtain the enhanced texture feature T. r
[0047]
[0048] in represents 1×1 convolution, represents a 3×3 dilated convolution with a dilation rate of 2. represents a 3×3 element-wise convolution, Represents 5×5 convolution;
[0049] Then, the texture of the thermal infrared image is enhanced in the vertical direction by using two similar branches. The upper branch splits the 3×3 convolution into a 3×1 convolution and a 1×3 convolution. The lower branch changes the order of the 3×1 convolution and the 1×3 convolution and adds a 1×1 convolution for dimensionality reduction. Finally, the upper and lower branches are connected in the channel dimension to obtain the thermal infrared image T. v :
[0050] T v =Cat(Conv(f 1×3 (f 3×1 (T r ))),Conv(f 3×1 (f 1×3 (T r )))) (19)
[0051] where f 1×3 , f 3×1 Represents 1×3 convolution, 3×1 convolution.
[0052] Compared with the prior art, the advantages of the present invention are:
[0053] 1. Improve image quality: Since the dual-branch edge enhancement method based on the Laplacian operator and the Sobel operator is used in step 2, the details and contours in the image can be highlighted at the same time. Traditional edge detection methods usually only perform well in specific types of images, and when facing thermal infrared images with low contrast or complex backgrounds, some details are often lost or noise is introduced; step 4 uses an adaptive contrast enhancement module. The present invention dynamically adjusts the contrast by calculating the standard deviation of each position, improves the performance of high-contrast areas, and suppresses the influence of low-contrast areas; the texture enhancement method in the horizontal and vertical directions of step 5 combines the dilated convolution, the element-by-element convolution and the ordinary convolution. The present invention can perform more comprehensive texture enhancement on thermal infrared images in the horizontal and vertical directions, and uses multiple convolution operations to extract texture features from different levels and scales.
[0054] 2. Good modality fusion and high precision: Since the present invention adopts the strategy of using the summed modality as the key and value of cross-attention and the other modality as the query for fusion, this design avoids the information loss that may be caused by direct interaction between modalities in traditional methods, and ensures that the specific information of each modality is better preserved. Through this interactive mode, the present invention can make more full use of the relationship between different modalities, promote the information interaction and synergy between enhanced modalities, give full play to the advantages of the two modalities, improve the quality of information fusion, and thus improve the overall performance.
[0055] 3. Ability to balance performance and accuracy: Since in the method of the present invention, step 2 adopts an asymmetric structure for enhancement in the visible light image enhancement module and the thermal infrared image enhancement module, the characteristics of the modality can be fully utilized, and the interactive fusion module of the modality can better utilize the complementary information of the two modalities, so that the tracker can adapt well to complex tracking scenarios.
[0056] 4. The present invention can realize efficient enhancement and fusion of visible light images and thermal infrared images, thereby making full use of the complementary information of the two modalities. Due to the use of an asymmetric structure for targeted enhancement, the characteristics of each modality are ensured to be optimally processed, thereby improving the tracking accuracy and robustness. The information of the two modalities is effectively integrated through an interactive fusion module, achieving more accurate target recognition and positioning. The present invention has strong robustness and adaptability when dealing with various complex tracking situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a principle block diagram for realizing the method of the present invention;
[0058] Figure 2 It is a process diagram of the enhancement module in the channel dimension of the present invention;
[0059] Figure 3 It is a detailed information enhancement diagram of the present invention;
[0060] Figure 4 It is the semantic information enhancement graph of the present invention;
[0061] Figure 5 The edge information-contrast enhancement map of the present invention;
[0062] Figure 6 are the texture enhancement maps in the horizontal and vertical directions of the present invention;
[0063] Figure 7 It is the interactive fusion module diagram of the present invention;
[0064] Figure 8 It is the accuracy graph and success rate graph of the tracking algorithm of the present invention on RGBT234;
[0065] Fig. 9 The performance evaluation results of different algorithms of the present invention on the RGBT210 dataset;
[0066] Fig.10 The performance evaluation results of different algorithms of the present invention on the LasHeR dataset. DETAILED DESCRIPTION
[0067] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0068] See also Figure 1 The tracking algorithm proposed in the present invention mainly includes four steps: feature extraction, feature enhancement, interactive fusion and prediction. In the feature extraction step, the backbone network is used to extract feature information from the visible light modality and the thermal infrared modality respectively. The feature information enters the feature enhancement step. In this step, the feature information is sent to a visible light image enhancement module and a thermal infrared image enhancement module in parallel for processing, and the enhanced visible light image and thermal infrared image are output; the enhanced visible light image and thermal infrared image enter the interactive fusion step, and the information interaction and fusion between the modalities are carried out in a Transformer-based interactive fusion module; finally, the prediction step is entered, and the prediction head is used for classification and regression.
[0069] The present invention proposes an RGBT target tracking method combining asymmetric enhancement and interactive fusion, and the specific steps are described as follows:
[0070] Step 1: Feature extraction: Use the backbone network to extract feature information from the visible light modality and thermal infrared modality respectively.
[0071] The proposed method uses two branches corresponding to a pair of images (visible light image and thermal infrared image) as the input of the backbone network in the feature extraction part. The present invention selects VGG-M as the backbone network for feature extraction. Specifically, the first three layers of the VGG-M network constitute the backbone extraction branch of the backbone network, which is used to extract the shared features of the visible light image and the thermal infrared image. The shared features are then added to the parallel backbone for further processing.
[0072] Step 2: Feature enhancement step: The feature information enters the feature enhancement step, in which the feature information is sent to a visible light image enhancement module and a thermal infrared image enhancement module in parallel for processing, and the enhanced visible light image and thermal infrared image are output. It includes two parallel parts: visible light image enhancement and thermal infrared image enhancement.
[0073] (I) Visible light image enhancement:
[0074] This section specifically proposes to perform detail-semantic enhancement on visible light images in the channel dimension, thereby improving the expressiveness of images by enhancing the detail information and semantic information of each channel.
[0075] The enhancement module process in the channel dimension is as follows Figure 2 As shown, considering that the detail information of the visible light image can better obtain the location information of the target, the semantic information of the visible light image can better obtain the category information of the target, the detail-semantic information of the visible light feature is enhanced in the channel dimension and the spatial dimension.
[0076] 1. Perform detail-semantic enhancement on the visible light image in the channel dimension. The information in each channel is weighted according to its importance in the entire image, which improves the expression ability of the local detail information and global semantic information of the image, and realizes detail-semantic enhancement in the channel dimension. The specific implementation method is as follows:
[0077] ε=S(δ(GAP(X)+GMP(X)) (1)
[0078] δ=f2(γ(f1)) (2)
[0079] X'=ε×X+(1-ε)×X (3)
[0080] Among them, f1 and f2 represent 1×1 convolution, γ represents the ReLU function, S represents the Sigmoid function, GAP and GMP represent global average pooling and global maximum pooling.
[0081] 2. Perform detail-semantic enhancement on visible light images in the spatial dimension. The detail-semantic enhancement in the spatial dimension captures information of different scales through multi-scale feature extraction, and then models the image context relationship through the spatial self-attention mechanism. At the same time as the spatial self-attention mechanism modeling, the spatial weighting mechanism is designed to enhance the spatial information of the image, thereby improving the expression ability of local details and global semantics in the image. The specific steps are as follows:
[0082] First, multi-scale feature extraction is performed on visible light images to cope with the challenges of scale changes and small-sized objects. Figure 3 As shown in Figure 1, multi-scale feature extraction includes a 1×1 convolution, three 3×3 dilated convolutions with different dilation rates, and global average pooling. Finally, the feature maps extracted from different scales are connected in the channel dimension to obtain the enhanced detail feature map F. The specific calculation is shown in Formula 4:
[0083] F=Cat(f3(X),f4(X),f5(X),f6(X),U(GAP(X)) (4)
[0084] Among them, Cat represents the connection operation, f3 represents 1×1 convolution, f4, f5, f6 represent dilation rates of 6, 12, and 18 respectively, U represents the upsampling operation, and F represents the enhanced detail features.
[0085] After local context information enhancement, the visible light detail feature is obtained, and the detail feature F is semantically enhanced, such as Figure 4 As shown, it contains two branches, namely the spatial self-attention branch and the spatial weighted branch.
[0086] Then the global context is modeled based on the spatial self-attention mechanism. Specifically, Q, K are modeled through the feature F and the corresponding weight matrix W. q and W k Obtained by linear transformation. The features after linear transformation are:
[0087] Q=W q F (5)
[0088] K=W k F (6)
[0089] Get a correlation matrix R.
[0090] Subsequently, the features obtained by modeling the spatial self-attention mechanism contain contextual dependencies on surrounding objects and have rich semantic information. The feature map R is globally max-pooled in the spatial dimension to obtain the global position encoding. After global max-pooling, the global semantic feature map R'∈R is obtained. 1×H×W , the global semantic feature map R' contains the relationship between different feature channels.
[0091] R=Q T K (7)
[0092] R'=S(Maxpool(Q T K) (8)
[0093] At the same time, through the feature F and the corresponding weight matrix W v Perform a linear transformation to obtain V∈R C×H×W :
[0094] V=W v F (9)
[0095] Take the dot product of the feature map V and R' to get the output V of the spatial self-attention module r :
[0096] V r =R'⊙V (10)
[0097] While modeling based on the spatial self-attention mechanism, a spatial weighting mechanism is designed to use the maximum function to highlight the significant features at each spatial position. By selecting the maximum value of the channel dimension at each position, it ensures that the features with high importance are given priority. The specific calculation formula is:
[0098] α=LRN(ReLU(Conv)) (11)
[0099] F1=α(max(F)) (12)
[0100] w=S(Conv(F1)) (13)
[0101] Then the output feature map V of the spatial attention branch is r Multiply it with the spatial weight w and add it element by element with the feature map F to get the final enhanced feature V. The mathematical expression is:
[0102]
[0103] (II) Thermal infrared image enhancement: The thermal infrared image enhancement module includes an edge information enhancement module, an adaptive contrast enhancement module and a texture information enhancement module. The edge information enhancement module is designed based on the Laplace operator and the Sobel operator. The thermal infrared features are first sent to the edge information enhancement module for enhancement, then enhanced by the adaptive contrast enhancement module, and then enhanced in the horizontal and vertical directions by the texture information enhancement module. The specific steps are described as follows:
[0104] 1. The thermal infrared image is enhanced by using edge enhancement modules based on Laplacian operator and Sobel operator to highlight the edge and contour of the target in the image, and the thermal infrared image contrast is enhanced by using an adaptive contrast enhancement module. The edge information enhancement module and the contrast enhancement module are as follows: Figure 5 shown.
[0105] The thermal infrared features T extracted by the backbone network h The edge information enhancement module adopts a double-branch structure. The upper branch is used for thermal infrared feature T h The Laplacian operator is used to highlight the edge features of the image by calculating the second-order derivative of the image. After that, a convolution block (Conv-BN-ReLU) is passed to finally add the edge feature maps obtained from the two branches element by element to obtain the final edge feature map T e :
[0106]
[0107] Where β=ReLU(BN(Conv)), La represents the Laplace operator, So represents the Sobel operator, and BN represents local response normalization.
[0108] 2. Thermal infrared image T after edge enhancement e The input is sent to the adaptive contrast enhancement module, which aims to improve the contrast of the thermal infrared image through adaptive contrast enhancement, making the boundaries between different temperature areas clearer and the target more prominent. The specific process of the adaptive contrast enhancement module is:
[0109] First, the standard deviation of each position (i, j) in the feature is calculated to measure the contrast feature of the position. The calculation formula is:
[0110]
[0111] Where w and h represent the width and height of the thermal infrared image, respectively, and μ ij , σ ij Respectively represent the mean and standard deviation of the image, and P(i,j) represents the pixel value of the i-th row and j-th column.
[0112] This is followed by two 3×3 convolutional layers with ReLU and Sigmoid activation functions to obtain the contrast weights of the feature maps, which help enhance high-contrast regions and suppress low-contrast regions.
[0113] Finally, the edge feature T e Multiplying with the contrast weights yields the enhanced feature T s :
[0114] T s =T e ×S(A(σ ij (T e )) (17)
[0115] Where A = Conv(ReLU(Conv)).
[0116] 3. After enhancing the edge information and contrast of the thermal infrared image, the information is sent to the texture information enhancement module to perform texture enhancement in the horizontal and vertical directions. Combining the texture enhancement strategies in the horizontal and vertical directions, through the synergistic enhancement in these two directions, the present invention can extract detail textures in different directions, thereby improving the overall quality of the image and the expression ability of texture information. Texture enhancement is as follows: Figure 6 shown.
[0117] Feature T sFirst, texture enhancement is performed in the horizontal direction of the thermal infrared image. Atrous convolution, element-by-element convolution and ordinary convolution can better extract the texture information of the image. The images obtained after different convolutions are connected and then a 1×1 convolution is performed to reduce the dimension to obtain the enhanced texture feature T. r :
[0118]
[0119] in represents 1×1 convolution, represents a 3×3 dilated convolution with a dilation rate of 2. represents a 3×3 element-wise convolution, Represents a 5×5 convolution.
[0120] Then, the texture of the thermal infrared image is enhanced in the vertical direction by using two similar branches. The upper branch splits the 3×3 convolution into a 3×1 convolution and a 1×3 convolution. The lower branch changes the order of the 3×1 convolution and the 1×3 convolution and adds a 1×1 convolution for dimensionality reduction. Finally, the upper and lower branches are connected in the channel dimension to obtain the thermal infrared image T. v :
[0121] T v =Cat(Conv(f 1×3 (f 3×1 (T r ))),Conv(f 3×1 (f 1×3 (T r )))) (19)
[0122] where f 1×3 , f 3×1 Represents 1×3 convolution, 3×1 convolution.
[0123] Step 3: Interactive integration:
[0124] This module uses two Transformer-based encoders to learn the specific features of the two modalities, and uses a Transformer-based decoder to interact with the two modalities and achieve information fusion between the modalities. Figure 7 shown.
[0125] The present invention proposes two parallel decoders for fusion. Different from other fusion methods, the present invention adds the two modes after self-enhancement and uses the added mode as the key K' and value V' of multi-head cross attention, while the other mode is fused as the query Q'.
[0126] The specific formula of Transformer's self-attention mechanism is as follows:
[0127]
[0128] in, is the key factor to normalize attention, Q, K, V are linear transformations derived from features. In this invention, we use a multi-head self-attention mechanism, and its operation process can be described as:
[0129] MHead(Q,K,V)=Cat(head1,…head n )W O (twenty one)
[0130]
[0131] in, Represents the weight matrix vectors of Q, K, and V, which are the same as Q, K, and V in self-attention. d model represents the sequence encoding length, N represents the number of heads of multi-head attention, where h = d model / N, so the two self-enhancement modules are expressed as:
[0132] V′ r =MHead(V r +P,V r +P,V r ) (twenty three)
[0133] T′ v =MHead(T v +P,T v +P,T v ) (twenty four)
[0134] Among them, P represents the position code, V' r and T' v They represent the outputs of the two modalities after passing through the multi-head self-attention module. In addition, feature enhancement will introduce noise information. In order to better fuse the two modalities, a soft threshold module is added after the self-enhancement module. The final outputs after the soft threshold module are V' and T' respectively:
[0135]
[0136] Among them, μ = FC (ReLU (FC)), R represents soft thresholding, then the two modalities after soft thresholding are added to obtain image U, and the added image U is used as the key K' and value V' of two multi-head cross attentions, and the visible light image and thermal infrared image after soft thresholding are used as the query Q' of multi-head cross attention. After interactive fusion, the final output is:
[0137] Output=Conv(Cat(MHead(T′+P,U+P,U),MHead(V′+P,U+P,U))) (27)
[0138] Among them, Cat represents the channel connection of the two modes after interactive fusion, Conv represents 1×1 convolution, which is used to reduce the channel dimension, and Output represents the final fusion feature, which will be used for subsequent positioning and tracking.
[0139] Step 4: Prediction:
[0140] The prediction head adopts the existing network tracker prediction head structure. The features obtained through the network are sent to three fully connected layers. The current target and background are identified by using the Softmax function to calculate the final target position.
[0141] Using the PyTorch deep learning framework, image pairs were extracted from four public large-scale datasets: GTOT, RGBT210, RGBT234, and LasHeR. During the training process, the convolutional layer parameters in the feature extraction were initialized using the pre-trained VGG-M model, and the parameters of other layers were randomly initialized. In addition, the learning rate of the convolutional layer was set to 0.0001, the training learning rate of the feature enhancement module and fusion module of the modality was 0.0001, and the training learning rate of the fully connected layer used for classification learning was set to 0.001. The weight decay and momentum were fixed to 0.0005 and 0.9, respectively.
[0142] In order to objectively evaluate the performance of the method of the present invention, the tracking algorithm is evaluated on four public authoritative benchmark datasets, GTOT, RGBT210, RGBT234 and LasHeR, together with 24 typical tracking algorithms, and the performance of the tracking algorithm is analyzed from multiple angles.
[0143] GTOT contains 50 pairs of visible and thermal infrared video sequences in different scenes and conditions, totaling about 15,800 frames. It consists of frames with manually labeled ground truth and divides the challenge attributes into seven groups according to the state of the target. The proposed RGBT tracker is able to achieve good performance with good accuracy (0.907) and success rate (0.753).
[0144] Table 1 Performance evaluation results of different algorithms on the GTOT dataset
[0145]
[0146] RGBT234: This dataset is a large-scale RGBT tracking dataset. It is an extended version of the RGBT210 dataset with 234 video sequences and 12 challenging attributes.
[0147] Depend on Figure 8 It can be seen that the performance of the tracking algorithm proposed in this invention is significantly better than that of several other algorithms such as SiamFC. Compared with the second-ranked network APFNeT, its success rate is increased by 2.6%. This shows that the proposed feature enhancement and interactive fusion method between modalities is effective.
[0148] Table 2 Attribute-based evaluation with other trackers on the RGBT234 dataset
[0149]
[0150]
[0151] It can be seen from Table 2 that the tracking algorithm proposed in this invention performs well in the challenges of NO, PO, DEF, FM, SV, CM and MB compared with other trackers, indicating that the proposed algorithm has strong robustness and adaptability in dealing with various complex tracking situations.
[0152] The RGBT210 dataset contains 210 videos with a total of about 210,000 frames. This algorithm is compared with other advanced RGBT trackers. The experimental results are as follows: Fig. 9 As shown in the figure, the accuracy and success rate of the RGBT tracker of this invention reach 82.0% and 58.8%, and it is worth noting that compared with the advanced tracker APFNet tracker, the algorithm improves the accuracy and success rate by 1.2% and 3.5%, respectively. This proves that the asymmetric enhancement structure and interactive fusion method proposed in this invention are effective.
[0153] LasHeR: This dataset consists of 1224 visible light and thermal infrared video pairs, with a total of more than 730K frame pairs. The comparison results on this dataset are as follows Fig.10 As shown in the figure, the accuracy and success rate reached 60.7% and 47.9%. It can be inferred that the enhancement of the asymmetric structure can make full use of the characteristics of the modality, and the interactive fusion module of the modality can better utilize the complementary information of the two modalities, so that the tracker can adapt well to complex tracking scenarios, further verifying the effectiveness of the algorithm.
[0154] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A RGBT target tracking method combining asymmetric enhancement and interactive fusion, characterized in that: It includes the following four steps: Step 1: Feature extraction: Use the backbone network to extract feature information from the visible light modality and thermal infrared modality respectively; Step 2, feature enhancement: The feature information enters the feature enhancement step, in which the feature information is sent to a visible light image enhancement module and a thermal infrared image enhancement module in parallel for processing, and two modes are output: an enhanced visible light image and a thermal infrared image; Step 3: Interactive fusion: Add the two modalities after self-enhancement, and use the added modality as the key K' and value V' of multi-head cross attention, and the other modality as the query Q' for fusion; Step 4: Prediction: Use the prediction head for classification and regression.
2. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 1, characterized in that: The visible light image enhancement module in the step 2 includes two parts: detail-semantic enhancement in the channel dimension and detail-semantic enhancement in the spatial dimension. The detail-semantic enhancement in the spatial dimension captures information of different scales through multi-scale feature extraction, and then models the image context relationship through a spatial self-attention mechanism. While modeling the spatial self-attention mechanism, a spatial weighting mechanism is used to design spatial weighting to enhance the spatial information of the image.
3. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 2, characterized in that: The multi-scale feature extraction includes one 1×1 convolution, three 3×3 dilated convolutions with different expansion rates and global average pooling. The feature maps extracted from different scales are connected in the channel dimension to obtain an enhanced detail feature map F, which is calculated as follows: F=Cat(f3X),f4(X),f5(X),f6(X),U(GAP(X)) (4) Among them, Cat represents the connection operation, f3 represents 1×1 convolution, f4, f5, f6 represent dilation rates of 6, 12, and 18 respectively, U represents the upsampling operation, and F represents the enhanced detail features.
4. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 2, characterized in that: The specific calculation formula based on the spatial self-attention mechanism modeling is: Q=W q F (5) K=W k F (6) R=Q T K (7) R'=S(Maxpool(Q T K) (8) Through the feature F and the corresponding weight matrix W v Perform a linear transformation to obtain V∈R C×H×W : V=W v F (9) Take the dot product of the feature map V and R' to get the output V of the spatial self-attention modeling module r The mathematical expression is: V r =R'⊙V(10)。 5. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 2, characterized in that: The specific calculation formula of the spatial weighting design of the spatial weighting mechanism is: α=LRN(ReLU(Conv)) (11) F1=α(max(F)) (12) w=S(Conv(F1)) (13) The mathematical expression of the final enhanced feature V is: 。 6. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to any one of claims 3 to 5, characterized in that: The thermal infrared image enhancement module in step 2 includes an edge information enhancement module, an adaptive contrast enhancement module and a texture information enhancement module. The edge information enhancement module is designed based on the Laplace operator and the Sobel operator. The thermal infrared features are first sent to the edge information enhancement module, and then passed through the adaptive contrast enhancement module. The enhanced information is then passed through the texture information enhancement module for texture enhancement in the horizontal and vertical directions respectively.
7. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 6, characterized in that: The edge information enhancement module adopts a double-branch structure, and the upper branch is used for thermal infrared feature T h The Laplacian operator is used to highlight the edge features of the image by calculating the second-order derivative of the image. After that, a convolution block (Conv-BN-ReLU) is passed to finally add the edge feature maps obtained from the two branches element by element to obtain the final edge feature map T e : Where β=ReLU(BN(Conv)), La represents the Laplace operator, So represents the Sobel operator, and BN represents local response normalization.
8. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 7, characterized in that: The specific process of the adaptive contrast enhancement module is as follows: First, input the edge-enhanced thermal infrared image feature T e , calculate the standard deviation of each position (i, j) in the feature, the calculation formula is: Where w and h represent the width and height of the thermal infrared image, respectively, and μ ij , σ ij Respectively represent the mean and standard deviation of the image, P(i,j) represents the pixel value of the i-th row and j-th column; Then there are two 3×3 convolutional layers, ReLU and Sigmoid activation functions to obtain the contrast weights of the feature map; Finally, the edge feature T e Multiplying with the contrast weights yields the enhanced feature T s : T s =T e ×S(A(σ ij (T e )) (17) Where A = Conv(ReLU(Conv)).
9. The RGBT target tracking method combining asymmetric enhancement and interactive fusion according to claim 8, characterized in that: The specific steps of the texture enhancement module are: Feature T s First, texture enhancement is performed in the horizontal direction of the thermal infrared image. Atrous convolution, element-by-element convolution and ordinary convolution are used to connect the images obtained after different convolutions, and then dimensionality reduction is performed to obtain the enhanced texture feature T. r : in represents 1×1 convolution, represents a 3×3 dilated convolution with a dilation rate of 2. represents a 3×3 element-wise convolution, Represents 5×5 convolution; Then, the texture of the thermal infrared image is enhanced in the vertical direction by using two similar branches. The upper branch splits the 3×3 convolution into a 3×1 convolution and a 1×3 convolution. The lower branch changes the order of the 3×1 convolution and the 1×3 convolution and adds a 1×1 convolution for dimensionality reduction. Finally, the upper and lower branches are connected in the channel dimension to obtain the thermal infrared image T. v : T v =Cat(Conv(f 1×3 (f 3×1 (T r ))),Conv(f 3×1 (f 1×3 (T r )))) (19) where f 1×3 , f 3×1 Represents 1×3 convolution, 3×1 convolution.
Citation Information
Cited By
Group behavior identification method and system based on cross-feature interaction Transform
CN120388335A
Visible light infrared target tracking method and system based on edge information dynamic enhancement
CN120807561A
Infrared unmanned aerial vehicle target tracking method based on morphological semantic collaborative perception
CN120808203A