An Unsupervised RGB-T Object Tracking Method Based on Attention Multi-Modal Feature Fusion

The proposed unsupervised RGB-T target tracking method integrates cross-level and cross-modal features using an attention-guided fusion module, addressing the limitations of existing supervised methods and enhancing tracking performance by reducing data dependency and costs.

CN114494354BActive Publication Date: 2025-07-15CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210138232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-07-15
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

Existing RGB-T target tracking algorithms face challenges in effectively fusing RGB and thermal infrared features due to the neglect of cross-level and cross-modal feature integration, and they rely heavily on supervised training methods that require significant labeled data, which is scarce in current datasets.

Method used

A no-supervised deep tracking model using a twin network architecture to align forward and backward tracking responses with initial labels, combined with an attention-guided feature fusion module to integrate features across levels and modalities, enhancing tracking performance.

Benefits of technology

The method leverages attention-driven feature fusion to exploit the strengths of different feature levels and modalities, reducing training costs and improving tracking accuracy through unsupervised learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494354B_ABST
    Figure CN114494354B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised RGB-T object tracking method based on attention multi-modal feature fusion. First, a hierarchical convolutional neural network is used to extract the features of RGB images and thermal infrared images. Then, a feature fusion module is used to synchronously fuse the features from different levels and different modalities. Next, the fused features are subjected to two forward tracking operations to obtain a response map. Then, the fused features are reversed, the original template image is used as the search image, the search image is used as the template image, and the generated response map is used as a pseudo-label for backward tracking to obtain the final response map. Then, the consistency loss between the response map obtained by backward tracking and the original label is minimized for unsupervised training. Finally, the test video frames are input into the trained network for forward tracking to obtain the response map, which is the predicted target position. The method of the present invention can make full use of multi-level and multi-modal information and can give play to the advantages of unsupervised learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an unsupervised RGB-T object tracking method based on attention multi-modal feature fusion, belonging to the RGB-T object tracking technology. Background Technique

[0002] Object tracking is an important research branch in the field of computer vision, and its goal is to predict the position of the object annotated in the first frame of the video. Although object tracking has made good progress recently, it still faces many challenges such as occlusion, low light, and large-scale changes.

[0003] Since the thermal infrared information is complementary to the RGB image under different environmental conditions, the tracking method that fuses RGB pictures and thermal infrared pictures is becoming more and more popular. How to better fuse the features of these two modalities for tracking is the key to the RGB-T object tracking algorithm. Existing algorithms usually use simple splicing or adding methods to fuse the information of the two modalities, and this method cannot make full use of the complementary information between the two modes. With the wide application of the attention mechanism in many fields of deep learning, some trackers use the attention mechanism to learn the weights of different modalities. However, when these methods perform fusion between different modalities, they ignore the feature fusion between different levels and cannot give full play to the advantages of features at different levels. High-level features can abstract more semantic features, while low-level features contain more texture detail features.

[0004] Existing RGB-T tracking algorithms generally use fully supervised training methods in the feature extraction process. Since this training method requires a large amount of labeled data, it will consume a certain amount of manpower and material resources. However, there are not many existing datasets for RGB-T object tracking. Therefore, with the increase in the amount of data, the unsupervised training method has great development potential in the RGB-T tracking task. Summary of the Invention

[0005] Object of the Invention: To solve the problems existing in the prior art, the present invention provides an unsupervised RGB-T object tracking method UDT-FF based on attention multi-modal feature fusion to reduce the training cost and improve the tracking performance of the RGB-T tracking algorithm. This method uses an unsupervised deep tracking model UDT as the backbone network, and uses a Siamese network architecture to compare the consistency between the response maps generated after forward tracking and backward tracking and the initial label to achieve unsupervised tracking; on the basis of the feature fusion AFF model, by constructing an attention-guided feature fusion FF module, the fusion between features at different levels and different modal features can be realized, and better tracking performance can be obtained.

[0006] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is:

[0007] An unsupervised RGB-T object tracking method based on attention multi-modal feature fusion, comprising the following steps:

[0008] (1) Input the RGB images of three video frames cropped to a size of 125×125 and their corresponding thermal infrared images into a two-layer convolutional network for feature extraction to obtain the first horizontal feature and the second horizontal feature of each image;

[0009] (2) Use a feature fusion module (abbreviated as the FF module) to fuse cross-modal and cross-level features simultaneously to obtain the fused template feature T and two search frame features S2 and S3;

[0010] (3) Perform forward tracking on the template feature T and the two search frame features S2 and S3 to obtain a forward tracking response map

[0011] (4) Swap the template feature T and the search frame feature S3, and use the forward tracking response map as pseudo-label supervision for backward tracking to obtain the final response map R;

[0012] (5) Calculate the consistency loss between the final response map R and the true label Q of the template feature T for unsupervised training to obtain a trained unsupervised RGB-T object tracking network;

[0013] (6) Input the video frame into the trained unsupervised RGB-T object tracking network for forward tracking to obtain a forward tracking response map, and the obtained forward tracking response map is the tracking result of the object.

[0014] Specifically, in step (1), the RGB images of three video frames and their corresponding thermal infrared images are input into a two-layer convolutional network for feature extraction. The RGB images of the three video frames are denoted as f1 v 、f2 v 、f3 v , and the corresponding thermal infrared images are denoted as f1 i 、f2 i 、f3 i ; is used as the template frame, and s2=(f2 v , f2 i ) and s3=(f3 v , f3 i ) are used as search frames; the first layer of the two-layer convolutional network is a filter with an output channel number of 3 and a convolution kernel size of 3×3. The first horizontal feature of the image is obtained through the first layer, f1 v 、f2 v 、f3v , f1 i , f2 i , f3 i The first-level features of a two-layer convolutional network The second layer is a filter with 32 output channels and a convolution kernel size of 3×3. The second-level features of the image, f1, are obtained through the second layer v , f2 v , f3 v , f1 i , f2 i , f3 i The second-level features of

[0015] Specifically, in step (2), a feature fusion module is used to fuse cross-modal and cross-level features simultaneously. The feature fusion module consists of three iterative AFF(a,b) functions, and the fusion process is as follows:

[0016]

[0017] Where: M(·) represents processing · using a multi-scale channel attention module (MS-CAM module, whose structural block diagram is as shown in Figure 2 ) represents pixel-wise addition, represents pixel-wise multiplication.

[0018] Specifically, in step (3), random labels are used as the true labels for unsupervised learning in the unlabeled video; in the training phase, the target template filter is used to perform forward tracking on the template feature T and the two search frame features S2 and S3. The target template filter is expressed as:

[0019]

[0020] Where: X represents the input of the target template filter, W X represents the output for the input X, Y represents the filtering label of X, and λ represents the regularization parameter; F(·) represents performing a discrete Fourier transform on ·, F -1 (·) represents performing an inverse discrete Fourier transform on ·, F * (·) is the conjugate complex number of F(·);

[0021] Performing forward tracking on the template feature T and the two search frame features S2 and S3 using the target template filter includes the following steps:

[0022] (31) Use the true label Q of the template feature T as the filtered label of the template feature T, and input the template feature T into the target template filter to obtain W T ;

[0023] (32) Calculate the response of the search feature S2 as:

[0024]

[0025] Where: is the response map generated by the template feature T and the search feature S2;

[0026] (33) Use as the filtered label of the search feature S2, input the search feature S2 into the target template filter to obtain

[0027] (34) Calculate the response of the search feature S3 as:

[0028]

[0029] Where: is the forward tracking response map.

[0030] Specifically, in step (4), swap the template feature T and the search frame feature S3, and use the forward tracking response map as pseudo-label supervision for backward tracking to obtain the final response map R, that is: use as the filtered label of the search frame feature S3, swap the template feature T and the search frame feature S3, and then use the method of step (3) to obtain the final response map R, including the following steps:

[0031] (41) Based on the search frame feature S3 and its filtered label calculate the response map of the search feature S2 as

[0032] (42) Use as the filtered label of the search feature S2, and calculate the response map of the template feature T based on the search feature S2 and its filtered label that is, the final response map R.

[0033] Specifically, in step (5), use the mean square error as the loss function to calculate the distance between the final response map R and the true label Q of the template feature T. The loss function is expressed as:

[0034] L = ||R - Q|| 2

[0035] Where: ||·|| 2Denote the sum of squared computed distances; theoretically, the final response map R should be consistent with the ground truth label Q, and the proposed unsupervised RGB-T object tracking method based on attention multi-modal feature fusion in this case is trained by minimizing this loss function.

[0036] Specifically, in step (6), the trained unsupervised RGB-T object tracking network is used to extract features from the RGB images and their corresponding thermal infrared images of two video frames, and the fused template feature T is obtained through the feature fusion module. * And the search frame feature S * , take the ground truth label Q * of the template feature T * as the filtering label, and calculate the response map of the search feature S * based on the template feature T * and its filtering label Q * , that is, the tracking result of the object.

[0037] Beneficial effects: The unsupervised RGB-T object tracking method based on attention multi-modal feature fusion provided by the present invention can give full play to the advantages of different-level and different-modal features by cross-level fusing RGB image and thermal infrared image features guided by attention; realize unsupervised object tracking by using temporal consistency, saving a large number of labeled RGB images and thermal infrared images, which is beneficial to the further development of unsupervised RGB-T object tracking. Brief Description of the Drawings

[0038] Figure 1 is the implementation flowchart of the method of the present invention;

[0039] Figure 2 is the structural block diagram of the MS-CAM module;

[0040] Figure 3 is the structural block diagram of the FF module for cross-modal and cross-level fusion;

[0041] Figure 4 is the structural schematic diagram of the device for implementing the method of the present invention. Detailed Embodiment

[0042] The present invention will be specifically introduced below in conjunction with the drawings and specific embodiments.

[0043] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0044] As Figure 1 shown is a flowchart of the implementation of an unsupervised RGB-T object tracking method based on attention multi-modal feature fusion, and the following will specifically describe each step.

[0045] Step S01: Input the RGB images of three video frames cropped to a size of 125×125 and their corresponding thermal infrared images into a two-layer convolutional network for feature extraction to obtain the first horizontal feature and the second horizontal feature of each image.

[0046] The unsupervised RGB-T object tracking network has a total of six images with a size of 125×125 as inputs, including three RGB images and their corresponding thermal infrared images. The three RGB images are denoted as f1 v , f2 v , f3 v , and the corresponding thermal infrared images are denoted as f1 i , f2 i , f3 i ; take f1 = (f1 v , f1 i ) as the template frame, and take s2 = (f2 v , f2 i ) and s3 = (f3 v , f3 i ) as the search frames; the first layer of the two-layer convolutional network is a filter with an output channel number of 3 and a convolutional kernel size of 3×3. Through the first layer, the first horizontal feature of the image is obtained. The first horizontal features of f1 v , f2 v , f3 v , f1 i , f2 i , f3 i are denoted as The second layer of the two-layer convolutional network is a filter with an output channel number of 32 and a size of 3×3. Through the second layer, the second horizontal feature of the image is obtained. f1v , f2 v , f3 v , f1 i , f2 i , f3 i The second-level features of

[0047] Step S02: Use the FF module to simultaneously fuse cross-modal and cross-level features to obtain the fused template feature T and two search frame features S2 and S3.

[0048] The structural diagram of the FF module is as Figure 3 shown. Use the FF module to simultaneously fuse cross-modal and cross-level features. The feature fusion module consists of three iterative AFF(a,b) functions, and the fusion process is as follows:

[0049]

[0050] Among them: M(·) represents processing · using the MS-CAM module; represents pixel-by-pixel addition, represents pixel-by-pixel multiplication.

[0051] The MS-CAM module refers to the multi-scale channel attention module, which is an attention method used in the FF module to generate channel attention weights. Its structural block diagram is as Figure 2 shown. After the MS-CAM module generates weights for the input features using the attention mechanism, it weights the input of size C×H×W. The MS-CAM module has two branches. One branch uses global average pooling to extract global information, and the other branch extracts local information. After using pointwise convolution to compress and restore the features of the channel dimension and then adding them together, the weights of the network are obtained.

[0052] Step S03: Perform forward tracking on the template feature T and the two search frame features S2 and S3 to obtain the forward tracking response map

[0053] In the unlabeled video, use random labels as the ground truth labels for unsupervised learning; in the training phase, use the target template filter to perform forward tracking on the template feature T and the two search frame features S2 and S3. The target template filter is expressed as:

[0054]

[0055] Among them: X represents the input of the target template filter, W X represents the output for the input X, Y represents the filtering label of X, λ represents the regularization parameter; F(·) represents performing the discrete Fourier transform on ·, F-1 (·) represents the inverse discrete Fourier transform of ·, F * (·) is the conjugate complex number of F(·);

[0056] Performing forward tracking on the template feature T and the two search frame features S2 and S3 using the target template filter, including the following steps:

[0057] (31) Taking the true label Q of the template feature T as the filtering label of the template feature T, and inputting the template feature T into the target template filter to obtain W T ;

[0058] (32) Calculating the response of the search feature S2 as:

[0059]

[0060] Where: is the response map generated by the template feature T and the search feature S2;

[0061] (33) Taking as the filtering label of the search feature S2, and inputting the search feature S2 into the target template filter to obtain

[0062] (34) Calculating the response of the search feature S3 as:

[0063]

[0064] Where: is the forward tracking response map.

[0065] Step S04: Swapping the template feature T and the search frame feature S3, and using the forward tracking response map as pseudo-label supervision for backward tracking to obtain the final response map R.

[0066] Taking as the filtering label of the search frame feature S3, after swapping the template feature T and the search frame feature S3, using the method of step (3) to obtain the final response map R, including the following steps:

[0067] (41) Calculating the response map of the search feature S2 based on the search frame feature S3 and its filtering label as

[0068] (42) Taking as the filtering label of the search feature S2, and calculating the response map of the template feature T based on the search feature S2 and its filtering label i.e., the final response map R.

[0069] Such as Figure 4As shown, steps 03 and 04 as a whole serve as the training process of the unsupervised RGB-T object tracking network.

[0070] Step S05: Calculate the consistency loss between the final response map R and the true label Q of the template feature T for unsupervised training to obtain the trained unsupervised RGB-T object tracking network.

[0071] Use the mean square error as the loss function to calculate the distance between the final response map R and the true label Q of the template feature T. The loss function is expressed as:

[0072] L = ||R - Q|| 2

[0073] where: ||·|| 2 represents calculating the sum of the squares of the distances to ·; theoretically, the final response map R should be consistent with the true label Q. By minimizing this loss function, the unsupervised RGB-T object tracking method proposed in this case is trained to obtain the next object tracking.

[0074] Step S06: Input the video frame into the trained unsupervised RGB-T object tracking network for forward tracking to obtain the forward tracking response map, and the obtained forward tracking response map is the tracking result of the object.

[0075] Use the trained unsupervised RGB-T object tracking network to extract features from the RGB images and their corresponding thermal infrared images of two video frames. After passing through the feature fusion module, the fused template feature T * and the search frame feature S * are obtained. Take the true label Q * of the template feature T * as the filtering label. Based on the template feature T * and its filtering label Q * calculate the response map of the search feature S * , which is the tracking result of the object.

[0076] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0077] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0078] The basic principles, main features, and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. An unsupervised RGB-T object tracking method based on attention multi-modal feature fusion, characterized in that: It includes the following steps: (1) Input the RGB images of three video frames resized to 125×125 and their corresponding thermal infrared images into a two-layer convolutional network Extract features to obtain the first horizontal features and the second horizontal features of each image; The RGB images of the three video frames are represented as The corresponding thermal infrared images are represented as Take as the template frame, and take and as the search frames; The first layer of the two-layer convolutional network is a filter with an output channel number of 3 and a convolutional kernel size of 3×3. The first horizontal feature of the image is obtained through the first layer, The first horizontal feature of The second layer of the two-layer convolutional network is a filter with an output channel number of 32 and a convolutional kernel size of 3×3. The second horizontal feature of the image is obtained through the second layer, The second horizontal feature of (2) Use the feature fusion module to simultaneously fuse cross-modal and cross-level features to obtain the fused template feature T and two search frame features S2 and S3; The feature fusion module consists of three iterative AFF(a,b) functions, and the fusion process is as follows: Wherein: M(·) represents processing · using a multi-scale channel attention module; represents pixel-wise addition, represents pixel-wise multiplication; (3) Forward tracking is performed on the template feature T and the two search frame features S2 and S3 using the target template filter to obtain a forward tracking response map. The target template filter is expressed as: Where: X represents the input of the target template filter, W X represents the output for the input X, Y represents the filtering label of X, and λ represents the regularization parameter; F(·) represents the discrete Fourier transform of ·, F -1 (·) represents the inverse discrete Fourier transform of ·, F * (·) is the conjugate complex number of F(); Use the target template filter to perform forward tracking on the template feature T and two search frame features S2 and S3, including the following steps: (31) Use the true label Q of the template feature T as the filtering label of the template feature T, and input the template feature T into the target template filter to obtain W T ; (32) Calculate the response of the search feature S2 as: Wherein: is the response map generated for the template feature T and the search feature S2; (33) Take as the filtering label of the search feature S2, and input the search feature S2 into the target template filter to obtain (34) Calculate the response of the search feature S3 as: Wherein: is a forward tracking response diagram; (4) Swap the template feature T and the search frame feature S3, and use the forward tracking response map as pseudo-label supervision to perform backward tracking to obtain the final response map R; (5) Calculate the consistency loss between the final response map R and the true label Q of the template feature T for unsupervised training to obtain a trained unsupervised RGB-T object tracking network; (6) Input the video frame into the trained unsupervised RGB-T object tracking network for forward tracking to obtain the forward tracking response map, and the obtained forward tracking response map is the tracking result of the target.

2. The unsupervised RGB-T object tracking method based on attention multi-modal feature fusion according to claim 1, characterized in that: In the said step (4), swap the template feature T and the search frame feature S3, and use the forward tracking response map as the pseudo-label supervision for backward tracking to obtain the final response map R, that is: use as the filtering label of the search frame feature S3. After swapping the template feature T and the search frame feature S3, adopt the method of step (3) to obtain the final response map R, including the following steps: (41)Based on the search frame feature S3 and its filtering label Calculate the response map of the search feature S2 as (42) Take as the filtering label of the search feature S2, and calculate the response map of the template feature T, that is, the final response map R, based on the search feature S2 and its filtering label ​ 3. The unsupervised RGB-T object tracking method based on attention multi-modal feature fusion according to claim 1, characterized in that: In the step (5), the mean square error is used as the loss function to calculate the distance between the final response map R and the true label Q of the template feature T, and the loss function is expressed as: L = ||R - Q|| 2 where: ||·|| 2 represents the sum of the squared distances calculated for · 4. The unsupervised RGB-T object tracking method based on attention multi-modal feature fusion according to claim 1, characterized in that: In the step (6), the trained unsupervised RGB-T object tracking network is used to extract features from the RGB images and their corresponding thermal infrared images of two video frames, and the fused template feature T is obtained through the feature fusion module. * and the search frame feature S * , the true label Q * of the template feature T * is used as the filtering label, and based on the template feature T * and its filtering label Q * the response map of the search feature S * is calculated, which is the tracking result of the target.

Citation Information

Patent Citations

  • RGBT target tracking method based on cross-modal attention mechanism and twin structure

    CN113628249A

  • Bimodal target tracking algorithm based on feature level and decision level fusion

    CN113920171A