RGBT tracking method based on CNN-Transform mixing

By combining CNN with Transformer, enhancing feature representation and combining weight generation networks for modal feature fusion, the problem that traditional convolutional neural networks are difficult to capture global information when processing long sequence data is solved, and the accuracy and robustness of RGBT target tracking is significantly improved.

CN120147357APending Publication Date: 2025-06-13XIAN UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510304742.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional convolutional neural networks are difficult to effectively capture global information when processing long sequence data, resulting in a degradation in target tracking performance in complex scenarios.

Method used

Using the RGBT tracking method based on CNN-Transformer hybrid, the template and search features are input into the Transformer module for enhancement, and combined with the weight generation network to fusion of modal features, and finally target prediction is performed through the box prediction head.

Benefits of technology

Significantly enhance feature representation capabilities, improve the accuracy and robustness of target tracking, especially when facing challenging scenario changes and target appearance changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147357A_ABST
    Figure CN120147357A_ABST
Patent Text Reader

Abstract

According to the RGBT tracking method based on CNN-Transform mixing, feature extraction is carried out on a template area by using an original AlexNet, and special extraction is carried out on a search area by using an improved AlexNet network; then results are input into a channel attention module for feature enhancement, template features and search features are obtained, the template features and the search features are spliced and then input into a Transform module, and branch features are obtained; the weight of each mode is generated through a weight generation network, and the branch features are fused according to the generated weights to obtain a fusion response graph # imgabs0 #; and performing target prediction on the fusion response graph # imgabs1 # through a frame prediction head to obtain a prediction bounding box bbox on the RGB image and the TIR image. According to the RGBT tracking method provided by the invention, the high efficiency of the convolutional network in local feature extraction and the strong capability of Transform in global information integration are fully utilized, the feature representation is remarkably enhanced, a complex tracking task can be efficiently and accurately completed, and the tracking accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and relates to an RGBT tracking method based on a CNN-Transformer hybrid. Background Art

[0002] The main content of object tracking is to predict key information such as the position and size of a target in subsequent frames given the information of the target in the initial frame of a certain video sequence. Object tracking has received extensive attention due to its great potential in many practical applications such as intelligent monitoring and autonomous driving systems. However, due to the imaging mechanism of visible light, some complex scenarios in open environments (such as illumination changes) limit the actual effect of object tracking based solely on RGB, resulting in target loss or incorrect tracking. Different from an RGB camera that captures the reflected light of an object, a thermal infrared (TIR) imaging sensor captures the heat emitted by the object itself. Compared with an RGB image that contains rich color textures under illumination conditions but fails in the dark, a TIR image provides significant contrast for hot objects, but has low resolution and poor texture. Therefore, in order to overcome the inherent disadvantages of single-modal methods, multi-modal object tracking has emerged, which makes full use of RGB and thermal images to perform more powerful all-weather tracking.

[0003] Existing RGBT trackers can be divided into two categories according to the type of features. One category is to use manually extracted features. However, since they cannot provide more discriminative feature information, their performance will decline when facing some challenging scenarios, such as cluttered backgrounds and occlusions. Inspired by the excellent performance of convolutional neural networks (CNNs) in the field of computer vision, many studies have used CNN features to improve tracking performance. RGBT trackers based on the CNN framework generally have obvious advantages over traditional trackers and can learn robust feature representations. However, due to the limitation of the receptive field, traditional convolutional neural networks (CNNs) are difficult to effectively capture global information when processing long-sequence data. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for RGBT tracking based on a CNN-Transformer hybrid, which solves the problem that traditional convolutional neural networks in the tracking method are difficult to effectively capture global information when processing long-sequence data due to the limitation of the receptive field.

[0005] The technical solution adopted by the present invention is a method for RGBT tracking based on a CNN-Transformer hybrid, and the steps are as follows: Step 1: Use the original AlexNet network as the feature extraction network for the template regions of visible light RGB and infrared TIR respectively; use the trained improved AlexNet network as the feature extraction network for the search regions of visible light RGB and infrared TIR respectively; obtain the template image z and the search image x , and input them into the above-mentioned feature extraction network respectively, and output the template region features of visible light RGB and the search region features , the template region features of infrared TIR and the search region features ; Step 2: Input and into the channel attention module for feature enhancement respectively to obtain the corresponding template features and search features ; Step 3: Concatenate the template features and the search features , and then input them into the Transformer module to obtain the visible light RGB branch features , the infrared TIR branch features ; Step 4: Input the branch features in Step 3 into the weight generation network to generate the weights for each modality, and then fuse , according to the generated weights to obtain the fusion response map ; Step 5: Perform object prediction on the fusion response map through the box prediction head to obtain the predicted bounding boxes bbox on the RGB image and the TIR image.

[0006] The feature of the present invention also lies in that Step 1 is specifically: use the original AlexNet network as the feature extraction network for the template regions of the visible light RGB branch and the feature extraction network for the template regions of the infrared TIR branch ; use the trained improved AlexNet network as the feature extraction network for the search regions of the visible light RGB branch and the feature extraction network for the search regions of the infrared TIR branch ; obtain the template image z and the search image x , and input them into the above-mentioned networks respectively, and output the visible light RGB search region features , the visible light RGB template region features , the infrared TIR search region features , the infrared TIR template region features 。

[0007] In step 1, the original AlexNet network was pre-trained on the ImageNet dataset, including 5 consecutively connected convolutional layers. After the first two convolutional layers, a max-pooling layer is connected respectively. After the first four convolutional layers, a dropout layer and a ReLU non-linear activation function are equipped. The trained improved AlexNet network includes 8 consecutively connected convolutional layers. A max-pooling layer is also connected between the third convolutional layer and the fourth convolutional layer, and a max-pooling layer is also connected between the fifth convolutional layer and the sixth convolutional layer. The convolutional kernel size of the first convolutional layer is 5×5. The fourth convolutional layer is a dilated convolutional layer with a convolutional kernel size of 3×3 and a stride of 2. The convolutional kernel sizes of the last three convolutional layers are all 3×3, and the strides are all 1.

[0008] In step 2, the channel attention module uses global average pooling (GAP) and a fully connected layer (FC) to dynamically assign weights to each channel for feature update and enhancement. The global average pooling operation compresses the information in the spatial dimension into a single value by calculating the global average of each feature channel. Based on the feature channel information compressed by the global average pooling operation, the fully connected layer learns and fits the complex mutual relationships between the feature channels to generate corresponding weights for each channel.

[0009] Specifically, step 2 is as follows: First, apply global average pooling (GAP) to the extracted features to obtain each compressed feature channel , and its calculation formula is: (1); where and represent the height and width of the feature respectively, is the feature of the th channel, and the weight coefficient indicating the reliability of the channel feature is: (2) where and represent two different fully connected layers, and are the sigmoid and ReLU activation functions respectively; finally, multiply the weight coefficient by the original feature to obtain the recalibrated feature: (3) where is the weighted th channel feature. is the weight coefficient of the confidence of the th channel feature; after passing through the channel attention module, the visible light RGB search feature , , and the search feature of the infrared TIR branch , and the template feature are obtained.

[0010] Step 3 is specifically as follows: Concatenate , , , to obtain the visible light RGB branch feature , the infrared TIR branch feature , and then input the concatenated feature into the Transformer module to obtain the visible light RGB branch feature , and the infrared TIR branch feature .

[0011] The Transformer module consists of an encoder layer, and the encoder layer consists of two sub-layer connection structures; The first sub-layer connection structure includes: a self-attention sub-layer (Self-Attention) for capturing long-range dependencies between features; a normalization layer (Layer Normalization) for stabilizing the training process and accelerating convergence; a residual connection (Residual Connection) for directly adding the input of the sub-layer to the normalization layer; The second sub-layer connection structure includes a feed-forward fully connected sub-layer (Feed-Forward Network, FFN) for performing non-linear transformation on features to further enhance the expression ability of features; a normalization layer (LayerNormalization); a residual connection (Residual Connection) for directly adding the input of the sub-layer to the normalization layer; The self-attention sub-layer adopts the self-attention mechanism, and the formula is as follows: (4) Among them, is 's dimensional parameter, and the input feature is multiplied by 's corresponding weight matrix to obtain the corresponding , and the specific value of will be continuously updated and learned; The input feature first passes through a multiplication by 's corresponding The weight matrix obtains the corresponding , and then and the matrix is multiplied to obtain the attention score . Then, dimensional scaling is performed, that is, the attention score is divided by the of the above formula. Finally, after passing through the activation function softmax , the obtained weight matrix is multiplied by V to obtain , and then the output of the first sub-layer is obtained through the normalization layer; The output of the first sub-layer is further processed through a feed-forward fully connected sub-layer, which consists of two fully connected layers with a ReLU activation function and a Dropout layer inserted in the middle: (5) The template features and search features of the visible light RGB and infrared TIR branches are respectively processed by the above formula (5), and then the output of the second sub-layer is obtained through the normalization layer and the residual connection, that is, the features of the visible light RGB branch and the infrared TIR branch are finally obtained respectively through the Transformer module and the features of the infrared TIR branch .

[0012] In step 4, the weight generation network consists of a global average pooling layer, ReLU, and Softmax activation functions. The specific calculation process is as follows: (6) (7) where GAP represents the global average pooling layer, represents the weight of the visible light RGB branch, represents the weight of the infrared TIR branch; The fused response map is calculated as follows: (8).

[0013] In step 5, the box prediction head is used to predict the probability distribution of the box corners, specifically: (9).

[0014] The beneficial effects of the present invention are: The method of RGBT tracking based on the CNN-Transformer hybrid of the present invention utilizes the self-attention mechanism in the Transformer architecture to achieve adaptive attention to useful semantic contexts, thereby significantly enhancing feature representation. This enhancement not only improves the expressiveness of template information within the relevant matching framework but also endows the network with sufficient modeling capabilities to efficiently and accurately complete complex tracking tasks. In this way, the network can more deeply understand and utilize the semantic information in the input data, thereby improving the accuracy and robustness of tracking, especially when facing challenging scene changes and target appearance changes; The method of RGBT tracking based on the CNN-Transformer hybrid of the present invention, a deep learning architecture that combines a convolutional network and a Transformer, can leverage the advantages of both in target tracking tasks. The convolutional network is responsible for extracting local features, while the Transformer enhances the expression and interaction of global features on this basis, enabling the generation of more discriminative response maps when extracting features from the search region and the template. This combination fully utilizes the efficiency of the convolutional network in local feature extraction and the powerful ability of the Transformer in global information integration, thereby achieving better performance in tasks such as RGBT target tracking. Through this fusion strategy, the model can not only accurately locate the target but also better adapt to the appearance changes of the target in different modalities and complex background interference, significantly improving the accuracy and robustness of tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a block diagram of the network structure of the RGBT tracking method based on the CNN-Transformer hybrid of the present invention; Figure 2 is an internal structure diagram of the Transformer module in the RGBT tracking method of the present invention; Figure 3 is a schematic diagram of the tracking result of the RGBT tracking method of Embodiment 6 of the present invention on an RGB image; Figure 4 is a schematic diagram of the tracking result of the RGBT tracking method of Embodiment 6 of the present invention on a TIR image. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Embodiment 1 The method of RGBT tracking based on the CNN-Transformer hybrid of the present invention has a process as Figure 1 shown, and the specific steps are as follows: Step 1: Use the original AlexNet network as the feature extraction network for the template region; use the trained improved AlexNet network as the feature extraction network for the search region; obtain the template image z and the search image x , with the image sizes as shown in Table 1 and Table 2 below. Input them into the above-mentioned feature extraction networks respectively, and output the template region feature Z and the search region feature X; Use the original AlexNet network as the feature extraction network for the visible light RGB branch template region and the infrared TIR branch template region ; use the trained improved AlexNet network as the feature extraction network for the visible light RGB branch search region and the infrared TIR branch search region , extract multi-layer features of the two modalities, where z is the template image, x denotes the search image, and its subscript represents the visible light branch and the infrared branch. Input them into the above-mentioned feature extraction networks respectively, and the output features after passing through the feature extraction network AlexNet are defined as the visible light RGB search region feature , the visible light RGB template region feature , the infrared TIR search region feature , the infrared TIR template region feature . A detailed introduction to the template and search region feature extraction network AlexNet is as follows.

[0018] For the feature extraction of the visible light RGB branch and the infrared TIR branch template regions, use the original AlexNet network, which is pre-trained on the ImageNet dataset, and its specific architecture is shown in Table 1. This network consists of 5 consecutive convolutional layers, with a max-pooling layer connected after each of the first two convolutional layers to reduce the feature dimension and enhance the invariance of the features. In addition, a dropout layer and a ReLU non-linear activation function are equipped after the first four convolutional layers. The dropout layer is used to prevent overfitting, while the ReLU activation function introduces non-linearity into the network, which helps the model learn complex feature representations. This structural design enables AlexNet to effectively extract rich features from image data, providing a solid foundation for subsequent tasks.

[0019] Table 1 Original AlexNet network structure parameters

[0020] In step 1, for the feature extraction of the visible light and infrared branch search regions, an improved AlexNet network is adopted. This network is a pre-trained model on the ImageNet dataset and has been subjected to specific optimization adjustments. Its detailed structure and parameter configuration are shown in Table 2. The improved version of the AlexNet network consists of 8 convolutional layers connected in sequence. The specific improvements include adding three new convolutional layers and placing the max pooling layer after the third and fifth convolutional layers to further reduce the spatial dimension of the features while enhancing the robustness of the features. Due to the different sizes of the template and the search region images, the size of the first convolutional layer changes from the original 11×11 to 5×5, and the stride remains 2 unchanged. The sizes of the three newly added convolutional layers are all 3×3, and the stride is 1. In addition, a key modification is made to the convolutional layer in the fourth stage of the network, changing the original 3×3 standard convolutional kernel to a dilated convolutional kernel with a stride of 2. This modification effectively expands the receptive field of the network, enabling the network to capture a larger range of image information, thereby helping to improve the model's comprehensive understanding ability of image features. Through these structural improvements, the improved AlexNet network demonstrates better performance and adaptability in processing image feature extraction tasks.

[0021] Table 2 Structural Parameters of the Improved AlexNet Network

[0022] In step 2, to further enhance the discriminative ability of the features, a channel attention module is introduced. The specific operation is to input the visible light RGB search region features , the visible light RGB template region features , the infrared TIR search region features , and the infrared TIR template region features into the channel attention module respectively for feature enhancement, obtaining the corresponding template features and search features of visible light RGB and infrared TIR.

[0023] The channel attention module can simulate the interdependent relationship between feature channels. By learning the weight distribution between different channels, it automatically highlights the more discriminative feature channels while suppressing the interference of noise, thereby enhancing the overall robustness of the features. This process not only helps the model better understand the key information in the data but also enables the subsequent analysis and decision-making processes to be more accurate and efficient. In this way, the channel attention network effectively improves the model's adaptability to complex scenarios and the fineness of feature expression.

[0024] The channel attention network uses global average pooling (GAP) and a fully connected layer (FC) to dynamically assign weights to each channel, thereby achieving feature update and enhancement. The global average pooling operation compresses the spatial dimension information into a single value by calculating the global average of each feature channel. This not only helps improve the convolutional neural network's perception ability of the target position but also reduces the number of parameters and avoids overfitting. Subsequently, based on this compressed feature channel information, the fully connected layer learns and fits the complex interrelationships between the feature channels, generating corresponding weights for each channel. These weights reflect the importance of different channels in the current task.

[0025] Specifically, first, apply global average pooling (GAP) to process the extracted features to obtain each compressed feature channel , and its calculation formula is: (1) where and represent the height and width of the feature respectively, is the feature of the th channel. Then, the weight coefficient indicating the channel feature credibility is: (2) where and represent two different fully connected layers, and are the sigmoid and ReLU activation functions respectively. Finally, multiply the weight coefficient by the original feature to obtain the recalibrated feature: (3) where is the weighted th channel feature, is the weight coefficient of the th channel feature credibility. The visible light RGB search feature , , and the search feature of the infrared TIR branch, and the template feature are obtained respectively. Each feature layer uses the channel attention module to focus on important regions, thereby improving the tracking performance.

[0026] Step 3: The search region and template features , , , Perform splicing to obtain the visible light RGB branch features , infrared TIR branch features , and then input the spliced features into the Transformer module to obtain the visible light RGB branch features , infrared TIR branch features . The detailed introduction of the Transformer module is as follows.

[0027] The Transformer module consists of an encoder layer, and the encoder layer consists of two sub-layer connection structures.

[0028] As Figure 2 shown, the first sub-layer connection structure includes a self-attention sub-layer (Self-Attention), which can capture the long-range dependencies between features. Following it is a normalization layer (LayerNormalization) for stabilizing the training process and accelerating convergence, and a residual connection (ResidualConnection), which directly adds the input of the sub-layer to the normalization layer, helping to solve the vanishing gradient problem in deep networks.

[0029] The second sub-layer connection structure includes a feed-forward fully connected sub-layer (Feed-Forward Network, FFN) for non-linearly transforming the features to further enhance the feature representation ability. Following it is also a normalization layer (Layer Normalization), and a residual connection (Residual Connection), which directly adds the input of the sub-layer to the normalization layer.

[0030] Through the processing of this encoder layer, the visible light RGB branch features , infrared TIR branch features are obtained. The template features and search region features of the two branches are subjected to feature extraction and feature enhancement. This process can not only integrate the information from the two modalities of visible light RGB and infrared TIR, but also learn richer and deeper feature representations through the advantages of the Transformer architecture, providing a higher-quality feature basis for subsequent tracking.

[0031] The self-attention mechanism is the core component of the self-attention sub-layer in the Transformer architecture. It allows the model to dynamically allocate different attention weights between different input positions, thereby capturing the long-range dependencies in the sequence data. The following are the specific implementation steps and related formulas of the self-attention mechanism: (4) Among them, is the dimensional parameter of. The input feature is multiplied by the corresponding weight matrix to obtain the corresponding , and the specific value of

[0032] will be continuously updated and learned. The input feature first passes through a multiplication by the corresponding weight matrix to obtain the corresponding , and then is multiplied by the matrix to obtain the attention score . Then, dimensionality scaling is performed (i.e., the attention score is divided by the softmax of the above formula). Finally, the weight matrix obtained after passing through the activation function is multiplied by V to obtain

[0033] The output of the first sub-layer is further processed through a feed-forward fully-connected sub-layer, which consists of two fully-connected layers with a ReLU activation function and a Dropout layer inserted in the middle: (5) The above formula (5) is respectively applied to the template features and search features of the visible light RGB and infrared TIR branches. After that, the output of the second sub-layer is obtained through a normalization layer and a residual connection, that is, the features of the visible light RGB and infrared TIR branches are finally obtained respectively through the Transformer module , .

[0034] Due to its locality characteristics, the convolution operation mainly focuses on the relationship between adjacent pixels when processing image data, and has limitations in capturing long-term dependencies. In contrast, as a global operation, the Transformer architecture can model the mutual relationship between all pixels in a single layer, thus more effectively capturing the long-term dependencies in the data. This global perception ability gives the Transformer significant advantages in processing complex patterns and long-distance information interaction.

[0035] The Transformer, through the self-attention mechanism, can directly calculate the relationship between any two positions in a sequence, thus better modeling the long-term dependencies of the target in the entire video sequence during object tracking. For example, when tracking a fast-moving target, the Transformer can utilize global information to better predict the future position of the target, even when the target is occluded or undergoes significant displacement changes in some frames.

[0036] With the aid of the self-attention mechanism in the Transformer architecture, the present invention realizes the adaptive attention to useful semantic contexts, thereby significantly enhancing the feature representation. This enhancement not only improves the expressiveness of the template information within the relevant matching framework but also endows the network with sufficient modeling capabilities to efficiently and accurately complete complex tracking tasks. In this way, the network can more deeply understand and utilize the semantic information in the input data, thereby improving the accuracy and robustness of tracking, especially when facing challenging scene changes and target appearance changes.

[0037] Step 4, in order to fully exploit and utilize the complementary advantages of visible light RGB and infrared TIR information, a fusion module is designed. This module can adaptively learn the weights of the RGB and TIR modal features, thereby achieving the effective fusion of the two modal responses. In the fusion module, the features of the visible light RGB and infrared TIR branches , are sent to the weight generation network to generate the weights for each modality. Then, according to the generated weights, the features of the visible light RGB and infrared TIR branches , are fused to obtain the final fused response map . The weight generation network consists of a global average pooling layer, Relu, and Softmax activation functions. The specific calculation process is as follows: (6) (7) where GAP represents the global average pooling layer, and represent the weights of the visible light RGB and infrared TIR branches respectively.

[0038] Finally, the calculation of the fused result, the fused response map is as follows: (8) Through this adaptive weight learning and fusion, this module can intelligently highlight the more informative modal features in a specific scenario while integrating the advantages of the two modalities, and finally generate a fused response map that combines RGB and TIR information Fusion response map It not only retains the key information of the two modalities, but also enhances the feature expression ability through the fusion strategy, providing more accurate and comprehensive support for subsequent predictions.

[0039] Step 5. Finally, the fused response map Performs object prediction through the box prediction head to obtain the predicted bounding boxes bbox on the RGB image and the TIR image. The detailed introduction of the box prediction head is as follows.

[0040] The RGB-T image pair is not always precisely aligned in each frame, and the position of the object in the infrared image may sometimes be different from its position in the visible light image. Therefore, it is difficult for the tracker to directly regress the coordinates of the object based on the RGB-T image pair. Thus, a box prediction head is used to predict the probability distribution of the box corners. The fused response map Is sent to the box prediction head to generate the predicted bounding box bbox. The calculation process is as follows: (9).

[0041] The deep learning architecture that combines the convolutional network and Transformer in the present invention can leverage the advantages of both in the object tracking task. The convolutional network is responsible for extracting local features, while the Transformer enhances the expression and interaction of global features on this basis, enabling the generation of more discriminative response maps when extracting features from the search region and the template. This combination fully utilizes the efficiency of the convolutional network in local feature extraction and the powerful ability of the Transformer in global information integration, thus achieving better performance in tasks such as RGBT object tracking. Through this fusion strategy, the model can not only accurately locate the object, but also better adapt to the appearance changes of the object in different modalities and complex background interference, significantly improving the accuracy and robustness of the tracking.

[0042] Embodiment 2 The RGBT tracking method based on the CNN-Transformer hybrid of the present invention is as follows: Step 1. Obtain the template image z and the search image x , and use the original AlexNet network to extract features from the visible light RGB and infrared TIR template regions respectively, and use the improved AlexNet network to extract features from the search region respectively, to obtain the template region features 、search region features 、template region features and search region features ; Step 2. Combine 、 They are respectively input into the channel attention module to obtain corresponding template features and search features ; Step 3: Concatenate the template features and search features and input them into the Transformer module to obtain visible light RGB branch features and infrared TIR branch features ; Step 4: Input the branch features in Step 3 into the weight generation network to generate weights for each modality, and fuse the branch features according to the weights to obtain a fused response map ; Step 5: Perform object prediction on the fused response map through the box prediction head to obtain the final predicted bounding box bbox

[0043] Embodiment 3 For the RGBT tracking method based on the CNN-Transformer hybrid of the present invention, on the basis of Embodiment 2, Step 1 is specifically: using the original AlexNet network as the feature extraction network for the template region of the visible light RGB branch and the feature extraction network for the template region of the infrared TIR branch ; using the trained improved AlexNet network as the feature extraction network for the search region of the visible light RGB branch and the feature extraction network for the search region of the infrared TIR branch ; obtaining the template image z and the search image x , inputting them into the above-mentioned networks respectively, and outputting the visible light RGB search region features , the visible light RGB template region features , the infrared TIR search region features , and the infrared TIR template region features .

[0044] The original AlexNet network was pre-trained on the ImageNet dataset, including 5 sequentially connected convolutional layers, with a max-pooling layer connected after each of the first two convolutional layers, and a random inactivation layer (Dropout) and a ReLU non-linear activation function were equipped after each of the first four convolutional layers; The trained improved AlexNet network includes 8 consecutively connected convolutional layers. There is also a max pooling layer connected between the third convolutional layer and the fourth convolutional layer, and a max pooling layer connected between the fifth convolutional layer and the sixth convolutional layer. The convolutional kernel size of the first convolutional layer is 5×5, the fourth convolutional layer is a dilated convolutional layer with a convolutional kernel size of 3×3 and a stride of 2, and the convolutional kernel sizes of the last three convolutional layers are all 3×3, and the strides are all 1.

[0045] Example 4 For the RGBT tracking method based on the CNN-Transformer hybrid of the present invention, on the basis of Example 3, step 2 is specifically as follows: First, apply global average pooling (GAP) to process the extracted features to obtain each compressed feature channel , and its calculation formula is: (1); where and respectively represent the height and width of the feature, is the feature of the th channel, and is the weight coefficient representing the channel feature credibility: (2) where and represent two different fully connected layers, and are the sigmoid and ReLU activation functions respectively; finally, multiply the weight coefficient by the original feature to obtain the recalibrated feature: (3) where is the weighted th channel feature, is the th channel feature credibility weight coefficient; through the channel attention module, the visible light RGB search feature , , and the search feature of the infrared TIR branch, and the template feature are obtained respectively.

[0046] Example 5 For the RGBT tracking method based on the CNN-Transformer hybrid of the present invention, on the basis of Example 4, step 3 is specifically as follows: Take , , , Perform splicing to obtain the visible light RGB branch features , and the infrared TIR branch features , and then input the spliced features into the Transformer module to obtain the visible light RGB branch features , and the infrared TIR branch features .

[0047] The Transformer module consists of an encoder layer, and the encoder layer consists of two sub-layer connection structures; The first sub-layer connection structure includes: a self-attention sub-layer (Self-Attention) for capturing long-range dependencies between features; a normalization layer (Layer Normalization) for stabilizing the training process and accelerating convergence; a residual connection (Residual Connection) for directly adding the input of the sub-layer to the normalization layer; The second sub-layer connection structure includes a feed-forward fully connected sub-layer (Feed-Forward Network, FFN) for performing non-linear transformation on features to further enhance the expression ability of features; a normalization layer (LayerNormalization); a residual connection (Residual Connection) for directly adding the input of the sub-layer to the normalization layer; The self-attention sub-layer adopts the self-attention mechanism, and the formula is as follows: (4) Among them, is the dimensional parameter of, the input feature is multiplied by the corresponding weight matrix to obtain the corresponding , the specific value of will be continuously updated and learned; The input feature first passes through a weight matrix corresponding to multiplying by to obtain the corresponding , and then and are multiplied to obtain the attention score , then dimensionality scaling is performed, that is, the attention score is divided by in the above formula, and finally the weight matrix obtained after passing through the activation function is multiplied by V to obtain softmax , and then the output of the first sub-layer is obtained through the normalization layer; After that, the output of the first sub-layer is obtained through the normalization layer; The output of the first sub - layer is further processed through a feed - forward fully - connected sub - layer, which consists of two fully - connected layers with a ReLU activation function and a Dropout layer inserted in the middle: (5) Process the template features and search features of the visible - light RGB and infrared TIR branches respectively using the above formula (5). Then, through a normalization layer and a residual connection, obtain the output of the second sub - layer, that is, finally obtain the features of the visible - light RGB branch and the infrared TIR branch respectively through the Transformer module and the features of the infrared TIR branch 。

[0048] Example 6 The RGBT tracking method based on CNN - Transformer hybrid of the present invention performs steps 1 to 5 of Example 1; The selected tracking video is the baby video sequence in the RGBT210 tracking dataset. Obtain the visible - light RGB search - area features ,visible - light RGB template features ,infrared TIR search - area features ,infrared TIR template features 。

[0049] Among them, in step 2, through the channel attention module, the search features and template features of the visible - light RGB and infrared TIR branches are obtained as follows: , , , 。

[0050] In step 3, the search features and template features of the RGB branch and the TIR branch obtained from step 2 , , , are concatenated to obtain the visible - light RGB branch features ,infrared TIR branch features ,and then the concatenated features are input into the Transformer module to obtain the visible - light RGB branch features ,infrared TIR branch features 。

[0051] In step 4, the features of the visible - light RGB and infrared TIR branches obtained from step 3 , are fused to obtain the final fused response map 。

[0052] In step 5, the fused response map obtained from step 4 The input is fed into the box prediction head for object prediction to obtain the predicted bounding boxes bbox on the RGB image and the TIR image.

[0053] The tracking result of a certain frame of the final baby video on the RGB image is as Figure 3 shown, and the tracking result on the TIR image is as Figure 4 shown. The baby video sequence contains 660 frames of pictures. Most frames of the video have strong light interference. There is strong light interference on the target object from frame 163 to 166 and from frame 482 to 586. However, between frame 92 and 105, there are challenges of illumination and motion blur. When the object moves away from these challenges, since the algorithm will rematch the initial template frame on the search image, it will be found again to achieve the tracking of the object. In fact, the target environment of this video sequence is not complex. However, in the case of strong light interference, it directly changes the appearance of the tracking target. And this algorithm actually has a powerful modeling ability, fully excavates the semantic information of the template features, and uses different neural networks to extract features of different depths, and can track the target under these challenges of appearance changes. The intersection over union (IOU) reflects the overlapping degree between the predicted tracking box and the manually annotated tracking box, and it is defined as the ratio of the area of the intersection of the predicted box and the ground truth box to the area of the union of the predicted box and the ground truth box.

[0054] (10) Among them, represents the predicted box, represents the ground truth box, represents the area. The larger the value of IOU, the higher the accuracy of the tracking algorithm, and its value range is . Usually, a success rate threshold is introduced, and it is required that when on a certain frame of picture, it is judged that the tracking is successful. Therefore, the accuracy metric is defined as the ratio of the number of images with on a video sequence to the total number of images in the video sequence.

[0055] For the baby video sequence. If the IOU threshold of the result is taken as , the accuracy of the baby video is 75%.

Claims

1. RGBT tracking method based on CNN-Transformer hybrid, characterized by: Here are the steps: Step 1: Get the template image z and the search image x The original AlexNet network is used to extract features from the visible light RGB and infrared TIR template areas respectively, and the improved AlexNet network is used to extract features from the search area respectively to obtain the template area features. , Search area features , Template area features and search area features ; Step 2: , Input them into the channel attention module respectively to obtain the corresponding template features and search features ; Step 3: Set the template feature and search features After splicing, it is input into the Transformer module to obtain the visible light RGB branch features , infrared TIR branch characteristics ; Step 4: Input the branch features of step 3 into the weight generation network to generate the weight of each modality, and fuse the branch features according to the weight to obtain the fusion response graph ; Step 5: Fusion response map The target is predicted through the box prediction head to obtain the final predicted bounding box bbox.

2. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: The step 1 is specifically as follows: the original AlexNet network is used as a feature extraction network for the visible light RGB branch template area and feature extraction network of infrared TIR branch template area ; The trained improved AlexNet network is used as the feature extraction network for the visible light RGB branch search area and the feature extraction network of the search area of ​​the infrared TIR branch ; Get template image z and search image x , input them into the above network respectively, and output the visible light RGB search area features , visible light RGB template area features , infrared TIR search area characteristics , infrared TIR template regional features .

3. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: In step 1, the original AlexNet network is pre-trained on the ImageNet dataset, including 5 convolutional layers connected in sequence, the first two convolutional layers are each followed by a maximum pooling layer, and the first four convolutional layers are each equipped with a random dropout layer (Dropout) and a ReLU nonlinear activation function; The trained improved AlexNet network includes 8 convolutional layers connected in sequence, a maximum pooling layer is connected between the third convolutional layer and the fourth convolutional layer, a maximum pooling layer is connected between the fifth convolutional layer and the sixth convolutional layer, the convolution kernel size of the first convolutional layer is 5×5, the fourth convolutional layer is an expanded convolution with a convolution kernel size of 3×3 and a step size of 2, and the convolution kernel sizes of the last three convolutional layers are all 3×3, and the step sizes are all 1.

4. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: In step 2, the channel attention module uses global average pooling GAP and fully connected layer FC to dynamically assign weights to each channel to update and strengthen features; The global average pooling operation compresses the information of the spatial dimension into a single value by calculating the global average value of each feature channel; The fully connected layer learns and fits the complex relationships between feature channels based on the feature channel information compressed by the global average pooling operation, and generates corresponding weights for each channel.

5. The RGBT tracking method based on CNN-Transformer hybrid according to claim 4, characterized in that: The step 2 is specifically as follows: firstly, the extracted features are processed by applying global average pooling GAP to obtain each compressed feature channel , and its calculation formula is: (1); in and Represent the height and width of the feature respectively, For the The characteristics of each channel represent the weight coefficient of the channel feature credibility. for: (2) in and represents two different fully connected layers, and are sigmoid and ReLU activation functions respectively; finally, the weight coefficient With the original features Multiply them together to get the recalibrated features: (3) in is the weighted Channel features, For the The weight coefficient of the credibility of the channel features; the visible light RGB search features are obtained through the channel attention module , , and the search features of the infrared TIR branch , template feature .

6. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: The step 3 is specifically as follows: Will , , , Splice to get the visible light RGB branch features , infrared TIR branch characteristics , and then input the spliced ​​features into the Transformer module to obtain the visible light RGB branch features , infrared TIR branch characteristics .

7. The RGBT tracking method based on CNN-Transformer hybrid according to claim 6, characterized in that: The Transformer module consists of an encoder layer, which consists of two sub-layer connection structures; The first sub-layer connection structure includes: a self-attention sub-layer (Self-Attention), which is used to capture the long-distance dependencies between features; a normalization layer (Layer Normalization), which is used to stabilize the training process and accelerate convergence; a residual connection (Residual Connection), which is used to add the input of the sub-layer directly to the normalization layer; The second sub-layer connection structure includes a feed-forward network (FFN) sub-layer, which is used to perform nonlinear transformation on features to further enhance the expressiveness of features; a normalization layer (LayerNormalization); a residual connection (Residual Connection), which is used to directly add the input of the sub-layer to the normalization layer; The self-attention sublayer adopts the self-attention mechanism, and the formula is as follows: (4) in, yes The dimension parameter of the input feature is multiplied by Corresponding The weight matrix can be obtained , The specific value of will be continuously updated and learned; The input features are first multiplied by Corresponding The weight matrix gets the corresponding ,Then and Matrix multiplication to get attention score , and then perform dimension scaling, that is, the attention score Divide by the above formula , and finally after the activation function softmax The obtained weight matrix is ​​multiplied by V to obtain , and then the output of the first sub-layer is obtained through the normalization layer; The output of the first sublayer is further processed through a feed-forward fully connected sublayer, which consists of two fully connected layers with a ReLU activation function and a Dropout layer inserted in between: (5) The template features and search features of the visible light RGB and infrared TIR branches are processed by the above formula (5), and then the output of the second sub-layer is obtained through the normalization layer and residual connection, that is, the features of the visible light RGB branch are finally obtained through the Transformer module. and the characteristics of the infrared TIR branch .

8. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: In step 4, the weight generation network consists of a global average pooling layer, Relu and Softmax activation functions. The specific calculation process is as follows: (6) (7) Where GAP represents the global average pooling layer, and Represent the weights of the visible light RGB and infrared TIR branches respectively; Fusion Response Plot The calculation of is as follows: (8)。 9. The RGBT tracking method based on CNN-Transformer hybrid according to claim 1, characterized in that: In step 5, the box prediction head is used to predict the probability distribution of the box corners, specifically: (9)。

Citation Information

Cited By

  • Transform-based cross-modal feature fusion and refining target tracking method

    CN120852473A

  • A target tracking method based on cross-modal feature fusion and refinement of transformer

    CN120852473B

  • RGB-T target tracking method and device

    CN121414788A