RGB-T Object Tracking Network Based on Convolution-Transformer with Global Attention

By adopting a convolution-Transformer network of global attention in RGB-T target tracking technology, combining the characteristics of convolution and Transformer, the global and local features of multimodal data are extracted, and the global correspondence between different modes is established through the cross attention structure, the problem of insufficient tracking performance in the existing technology is solved, and more efficient information interaction and fusion is achieved, and tracking performance and speed are improved.

CN115375948BActive Publication Date: 2025-06-03CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211113015.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-14
Publication Date
2025-06-03
Estimated Expiration
2042-09-14

AI Technical Summary

Technical Problem

The existing RGB-T target tracking technology has insufficient tracking performance when facing environmental interference such as occlusion, low light and fast movement, and has failed to fully utilize the complementary advantages of RGB and infrared information.

Method used

The convolution-Transformer network based on global attention is adopted, and the real-time multi-domain target tracking model RT-MDNet is transformed into a dual-stream multi-branch network. Combined with the characteristics of convolution and Transformer, the global and local features of multimodal data are extracted, and the global correspondence between different modes is established through the cross attention structure to achieve deep interaction and fusion.

Benefits of technology

The tracking performance and speed of the RGB-T target tracking algorithm are improved, and multimodal data can be used more effectively, and information interaction and fusion can be enhanced, thereby achieving more robust tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375948B_ABST
    Figure CN115375948B_ABST
Patent Text Reader

Abstract

The present invention discloses an RGB-T object tracking network method based on global attention convolutional-Transformer. The input is a pair of visible light and infrared images from a video stream. By constructing a convolutional-Transformer feature extraction network, the shared-parameter Transformer branch processes the shared features of multi-modal data and establishes a global information model. The independent convolutional branch processes the independent features of each modality. The global and local, shared and independent information features are combined to fully exploit the available information of multi-modal data. Then, through the cross-attention fusion module, cross-attention is used to establish the global correspondence between modalities to strengthen the information interaction between modalities and achieve deep fusion. The candidate samples are mapped into features of a fixed size through the ROIAlign layer. Finally, the fused features mapped by each candidate sample are fed into the fully connected layer to predict the scores obtained by the binary classification layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and is a method for RGB-T object tracking network based on global attention convolution-Transformer, belonging to the multi-modal RGB-T object tracking technology. Background Art

[0002] Object tracking, as an important task in the field of computer vision, aims to predict the position of an object in subsequent frames given the position coordinates of the object in the first frame of a video. However, the tracking task usually faces many challenges of external environmental interference, such as occlusion, low light, fast movement, etc.

[0003] RGB-T tracking is an important research direction in the field of object tracking in recent years, and can still achieve good tracking performance under the above challenges. The main reason is that it simultaneously utilizes the advantages of RGB images and infrared images. Due to different imaging methods and principles, RGB images and infrared images have different characteristics. RGB images have high resolution and clear texture, and can provide complete visual information in normal environments, while infrared images have low resolution and single information, but can provide supplementary information for RGB images in extremely dark and other harsh environments. Therefore, the main purpose of RGB-T tracking is to utilize the complementary advantages of RGB images and infrared images to overcome environmental interference and obtain more feature representations to achieve robust tracking of videos. In recent years, many people have been committed to promoting the progress of RGB-T tracking, but there is still a large research space due to the insufficient utilization of RGB and thermal infrared information.

[0004] Transformer is a structure that has emerged in recent years with the attention mechanism as the core. Different from the characteristics of convolutional networks that are dominant in local areas, Transformer has excellent global information modeling capabilities. Summary of the Invention

[0005] Object of the Invention: In order to improve the tracking performance of RGB-T tracking algorithms, the present invention provides a convolutional-Transformer RGB-T object tracking network based on global attention to fully exploit modal information; this method uses the real-time multi-domain object tracking model RT-MDNet as the backbone network, transforms it into a two-stream multi-branch network to simultaneously process multi-modal data, and simultaneously utilizes the different characteristics of Transformer and convolution to fully utilize global and local information; based on the cross-attention structure, this method simultaneously performs operations on two modalities and retrograde cross-attention to construct global correspondences between different modalities, realizes deep interaction and fusion, and can obtain better tracking performance.

[0006] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is:

[0007] A convolutional-Transformer RGB-T object tracking network method based on global attention, comprising the following steps:

[0008] (1) Obtain a training sample set from the RGB-T data stream, where each training sample includes a corresponding RGB image R and an infrared image T;

[0009] (2) Construct a convolutional-Transformer network to stepwise extract multi-modal image features of the RGB image R and the infrared image T, and aggregate the convolutional features and Transformer features to obtain the RGB feature F vis and the infrared feature F inf ;

[0010] (3) Construct a cross-attention fusion module to fuse the RGB feature F vis and the infrared feature F inf to strengthen the information interaction between the RGB image R and the infrared image T, and obtain the fused feature F out ;

[0011] (4) Map the candidate samples to be predicted onto the fused feature, and perform binary classification prediction through a fully connected layer. The candidate samples to be predicted belong to a certain frame of picture or a part of a certain frame of picture in the RGB-T data stream;

[0012] (5) Use the binary classification loss and the embedding loss to supervise the training of the binary classification prediction, and obtain a trained convolutional-Transformer tracking network;

[0013] (6) Input the video stream and the image to be tracked into the trained convolutional-Transformer tracking network, and track the image to be tracked in the video stream. The image to be tracked belongs to a certain frame of picture or a part of a certain frame of picture in the video stream.

[0014] Specifically, in step (2), the convolutional-Transformer network is used to stepwise extract the multi-modal image features of the RGB image R and the infrared image T in the training sample, including the following steps:

[0015] (21) The convolutional branch network layer-by-layer extracts features from the input image through three convolutional blocks, and sequentially obtains the first, second, and third layer modal features f 1 , f 2 , f 3 . Take f 3 as the convolutional feature; the RGB image R passes through the convolutional branch network to obtain f 1 vis , the infrared image T passes through the convolutional branch network to obtain f1 inf and

[0016] (22) The Transformer branch network performs input processing on the input image to obtain X 0 , and then performs feature extraction on the input image layer by layer through four groups of Transformer modules, successively obtaining the first, second, third, and fourth layer global shared features X 1 , X 2 , X 3 , X 4 , and finally performs output processing to obtain X, taking X as the Transformer feature; the RGB image R passes through the Transformer branch network to obtain and X vis , and the infrared image T passes through the Transformer branch network to obtain and X inf ;

[0017] (23) Element-wise addition and X vis , and then obtains the RGB feature F after layer normalization operation vis ; Element-wise addition and X inf , and then obtains the infrared feature F after layer normalization operation inf .

[0018] Specifically, in the step (21), the convolutional branch network consists of two convolutional branches with non-shared parameters, where one convolutional branch is used to extract the features of the RGB image R, and the other convolutional branch is used to process the features of the infrared image T; each convolutional branch consists of three sequentially connected convolutional blocks;

[0019] The first convolutional block includes a filter, a ReLU activation function, a local response normalization layer, and a max pooling layer, and is used to obtain the first layer modal features. The filter is a filter with an output channel number of 96 and a convolutional kernel size of 7×7;

[0020] The second convolutional block includes a filter, a ReLU activation function, a local response normalization layer, and a max pooling layer, and is used to obtain the second layer modal features. The filter is a filter with an output channel number of 256 and a convolutional kernel size of 5×5;

[0021] The third convolutional block includes a filter and a ReLU activation function, and is used to obtain the third layer modal features. The filter is a filter with an output channel number of 512 and a convolutional kernel size of 3×3.

[0022] Specifically, in step (22), the Transformer branch network obtains the globally shared features according to the following method:

[0023] (221) Perform input processing on the input image with an unfixed size, change the number of channels of the input image from 3 to 64, and obtain X 0 ;

[0024] (222) Complete the extraction of globally shared features through four stages. In each stage, use a group of Transformer modules to process the input features; in the first, second, and third stages, first use the Transformer module group to process the input features, and then downsample the output of the Transformer module group. The result of downsampling is used as the input feature of the next Transformer module group;

[0025] In each Transformer module, the input features are first flattened, then formed into a feature vector I through layer normalization operations, and then linear mappings are performed through different linear transformation layers to obtain the query set Q, the key set K, and the value set V. Then, self-attention operations are performed on the query set Q, the key set K, and the value set V. Next, the input features and the result of the self-attention operation are connected through a residual connection, and finally, the output of the Transformer module is formed through layer normalization operations; expressed as:

[0026] Q = L Q (I), K = L K (I), V = L V (I)

[0027] Attention(Q, K, V) = softmamax(QK T )V

[0028] SA(I) = Attention(Q, K, V)

[0029] where: L Q , L K , L V respectively represent the linear transformation layers of the query set, the key set, and the value set in the Transformer module; Attention represents the operation function of the self-attention mechanism, softmamax represents the normalized exponential function for multi-classification, and SA represents the process function of self-attention in the Transformer module;

[0030] (223) The outputs of the first, second, third, and fourth stages are X 1 , X 2 , X 3 , X 4 , for X 4Perform output processing to obtain X, that is, use a pooling layer to process X 4 to be the same size as f 3 X of the same size, achieving alignment between features, and enabling element-wise aggregation operations in subsequent operations.

[0031] Specifically, in step (222), the extraction of globally shared features is completed through four stages. In the first, second, third, and fourth stages, the number of Transformer modules in the Transformer module group is 2, 2, 2, and 1 respectively, and the Transformer modules in the same Transformer module group are connected in sequence.

[0032] Specifically, in step (222), the extraction of globally shared features is completed through four stages. In the first, second, and third stages, first use the Transformer module group to process the input features, and then downsample the output of the Transformer module group. Through downsampling, the feature sizes of the outputs of the Transformer module groups in the first, second, and third stages are adjusted to H, W, and C respectively represent the length, width, and number of channels of X 0 respectively.

[0033] Specifically, in step (3), a cross-attention fusion module is constructed to strengthen the information interaction and fusion between the RGB image R and the infrared image T, including the following steps:

[0034] (31) Obtain the key set of the RGB feature F vis through the linear transformation layer of the key set and value set and the value Obtain the query set of the infrared feature F inf through the linear transformation layer of the query set Perform cross-attention operation to obtain the cross feature

[0035]

[0036] CA(F vis , F inf ) = Attention(Q inf , K vis , V vis )

[0037] where: CA represents the operation function of cross-attention, and LN represents the layer normalization operation; respectively represent the linear transformation layers of the query set, key set, and value set in the cross-attention branch for calculating the cross feature ;

[0038] (32) Obtain the infrared feature F through the linear transformation layer of the key set and value set inf of the key set and the value set Obtain the RGB feature F through the linear transformation layer of the query set vis of the query set Perform cross-attention operation to obtain cross features

[0039]

[0040] CA(F inf , F vis ) = Attention(Q vis , K inf , V inf )

[0041] Where: respectively represent the linear transformation layers of the query set, key set and value set in the cross-attention branch for calculating cross features ;

[0042] (33) Perform a fusion operation to obtain the fusion feature F out :

[0043]

[0044] Where: Concat represents the operation function for feature fusion.

[0045] Specifically, in step (4), first map the candidate samples to be predicted into a set size (3×3) on the fusion feature F out , and then predict the binary classification scores of the candidate samples through a continuous set of fully connected layers.

[0046] Specifically, the continuous set of fully connected layers includes two general fully connected layers (FC4, FC5) and one multi-domain fully connected layer (FC6). The multi-domain fully connected layer is composed of a set of parallel binary classification layers, and the binary classification scores of the candidate pictures are obtained through the binary classification layers.

[0047] Specifically, in step (5), binary classification loss and embedding loss are used to supervise and train the binary classification prediction:

[0048]

[0049]

[0050] L = L cls + λL inst

[0051] Where: Lcls represents the binary classification loss; L inst represents the embedding loss, used to overcome the situation where candidate samples are different from but overly similar to certain classification categories; L represents the total loss function, and λ represents the weight;

[0052] i = 1, 2, …, N, where N represents the total number of candidate samples; c ∈ {0, 1}, c = 1 indicates that the true label of the candidate sample is positive, and c = 0 indicates that the true label of the candidate sample is negative; y i represents the probability that the system correctly judges the true label of the i-th candidate sample; f i d represents the binary classification score of the i-th candidate sample on classification d, d = 1, 2, …, D, where D represents the total number of categories; [·] cd represents normalizing · for both c = 1 and c = 0 cases for classification d; [·] d represents normalizing · for c = 1 case for classification d; σ represents the softmax function.

[0053] Beneficial effects: The RGB-T object tracking network method based on global attention convolutional-Transformer provided by the present invention combines global and local, shared and independent information features, and can fully exploit the available information in multi-modal data; at the same time, the cross-fusion module uses cross-attention to establish global correspondence relationships between modalities to strengthen information interaction, enabling deep fusion and improving tracking performance and speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is the implementation flowchart of the method of the present invention;

[0055] Figure 2 is the structural schematic diagram of the system of the present invention;

[0056] Figure 3 is the structural block diagram of the first, second, and third Transformer module groups;

[0057] Figure 4 is the structural block diagram of the fourth Transformer module group;

[0058] Figure 5 is the structural block diagram of the cross-fusion module. DETAILED DESCRIPTION OF THE INVENTION

[0059] The present invention will be specifically introduced below in conjunction with the accompanying drawings and specific embodiments.

[0060] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.

[0061] As Figure 1 shown is a RGB-T object tracking network method based on global attention convolutional-Transformer, and the following will specifically describe each step.

[0062] Step S01: Obtain a training sample set from the RGB-T data stream, and each training sample includes a corresponding RGB image R and an infrared image T.

[0063] Step S02: Construct a convolutional-Transformer network to stepwise extract multi-modal image features of the RGB image R and the infrared image T, and aggregate the convolutional features and Transformer features to obtain the RGB feature F vis and the infrared feature F inf .

[0064] (21) The convolutional branch network layer-by-layer extracts features of the input image through three convolutional blocks, and successively obtains the first, second, and third layer modal features f 1 , f 2 , f 3 , and takes f 3 as the convolutional feature; the RGB image R passes through the convolutional branch network to obtain f 1 vis , the infrared image T passes through the convolutional branch network to obtain f 1 inf ,

[0065] As Figure 2 shown, the convolutional branch network consists of two convolutional branches with non-shared parameters, one convolutional branch is used to extract the features of the RGB image R, and the other convolutional branch is used to process the features of the infrared image T; each convolutional branch consists of three successively connected convolutional blocks; the first convolutional block includes a filter, a ReLU activation function, a local response normalization layer (LRN), and a max pooling layer, and is used to obtain the first layer modal feature f 1, the filter is a filter with 96 output channels and a convolution kernel size of 7×7; the second convolutional block includes a filter, a ReLU activation function, a local response normalization layer (LRN), and a max pooling layer for obtaining the second-layer modal feature f 2 , the filter is a filter with 256 output channels and a convolution kernel size of 5×5; the third convolutional block includes a filter and a ReLU activation function for obtaining the third-layer modal feature f 3 , the filter is a filter with 512 output channels and a convolution kernel size of 3×3.

[0066] (22) The Transformer branch network performs input processing on the input image to obtain X 0 , and then uses four groups of Transformer modules to extract features from the input image layer by layer, obtaining the first, second, third, and fourth-layer global shared features X 1 , X 2 , X 3 , X 4 , and finally performs output processing to obtain X, taking X as the Transformer feature; the RGB image R passes through the Transformer branch network to obtain and X vis , the infrared image T passes through the Transformer branch network to obtain and X inf .

[0067] As Figure 2 shown, the Transformer branch network consists of two Transformer branches with shared parameters, where one convolutional branch is used to extract the features of the RGB image R, and the other convolutional branch is used to process the features of the infrared image T; the specific operation process of the Transformer branch is as follows:

[0068] (221) Perform input processing on the input image of H×W×3, changing the number of channels of the input image from 3 to 64 to obtain X 0 .

[0069] (222) Complete the extraction of global shared features through four stages. In each stage, a group of Transformer modules is used to process the input features.

[0070] As Figure 3As shown, in the first, second, and third stages, the Transformer module group includes two sequentially connected Transformer modules. First, the input features are processed by the Transformer module group, and then the output of the Transformer module group is downsampled. The result of the downsampling serves as the input features for the next Transformer module group.

[0071] As Figure 4 shown, in the fourth stage, the Transformer module group includes only one Transformer module.

[0072] In each Transformer module, the input features are first flattened, then passed through layer normalization operations to form the feature vector I. Then, different linear transformation layers are used to perform linear mapping to obtain the query set Q, key set K, and value set V. Next, self-attention operations are performed on the query set Q, key set K, and value set V. Then, the input features and the result of the self-attention operations are connected through a residual connection. Finally, layer normalization operations are performed again to form the output of the Transformer module. The processing process of the Transformer module can be expressed as:

[0073] Q = L Q (I), K = L K (I), V = L V (I)

[0074] Attention(Q, K, V) = softmax(QK T )V

[0075] SA(I) = Attention(Q, K, V)

[0076] where: L Q , L K , L V respectively represent the linear transformation layers for the query set, key set, and value set in the Transformer module; Attention represents the operation function of the self-attention mechanism, softmax represents the normalized exponential function for multi-classification, and SA represents the process function of self-attention in the Transformer module.

[0077] In the first, second, and third stages, first, the input features are processed by the Transformer module group, and then the output of the Transformer module group is downsampled. Through downsampling, the feature sizes of the outputs of the Transformer module groups in the first, second, and third stages are respectively adjusted to H, W, and C respectively represent the length, width, and number of channels of X 0 .

[0078] (223) The outputs of the first, second, third, and fourth stages are X 1 , X 2 , X 3 , X 4 , respectively. Output processing is performed on X 4 to obtain X, that is, the pooling layer is used to change X 4 to X with the same size as f 3 , realizing the alignment between features.

[0079] In step (22), the specific operation process for the RGB image R and the infrared image T can be described as follows:

[0080]

[0081]

[0082] Among them: LN represents layer normalization operation, Downsample represents downsampling operation, the superscript vis represents the operation for the RGB image R, the superscript inf represents the operation for the infrared image T, and the subscript i_j represents the output feature of the jth Transformer module in the ith stage.

[0083] (23) Element-wise addition and X vis , and after layer normalization operation, the RGB feature F vis is obtained; element-wise addition and X inf , and after layer normalization operation, the infrared feature F inf is obtained.

[0084] Step S03: Construct a cross-attention fusion module to fuse the RGB feature F vis and the infrared feature F inf , strengthen the information interaction between the RGB image R and the infrared image T, and obtain the fused feature F out .

[0085] As Figure 5 shown, the fusion process includes the following steps:

[0086] (31) Obtain the key set vis of the RGB feature F and the value set through the linear transformation layer of the key set and value set; obtain the query set inf of the infrared feature F through the linear transformation layer of the query set; perform cross-attention operation to obtain the cross feature

[0087]

[0088] CA(F vis ,F inf ) = Attention(Q inf ,K vis ,V vis )

[0089] Among them: CA represents the operation function of cross-attention, and LN represents the layer normalization operation; respectively represent the linear transformation layers of the query set, key set, and value set in the cross-attention branch for calculating cross features .

[0090] (32) Obtain the key set inf of the infrared feature F and the value set through the linear transformation layers of the key set and value set, and obtain the query set vis of the RGB feature F through the linear transformation layer of the query set, and perform cross-attention operation to obtain cross features

[0091]

[0092] CA(F inf ,F vis ) = Attention(Q vis ,K inf ,V inf )

[0093] Among them: respectively represent the linear transformation layers of the query set, key set, and value set in the cross-attention branch for calculating cross features .

[0094] (33) Perform a fusion operation to obtain the fused feature F out :

[0095]

[0096] Among them: Concat represents the operation function of feature fusion.

[0097] Step S04: Map the candidate sample to be predicted on the fused feature, and perform binary classification prediction through the fully connected layer. The candidate sample to be predicted belongs to a certain frame of picture or a part of a certain frame of picture in the RGB-T data stream.

[0098] First, pass the candidate sample to be predicted through the ROIAlign layer on the fused feature F outThe features are mapped into a 3×3 size, and then the binary classification scores of candidate samples are predicted through a continuous set of fully connected layers; the continuous set of fully connected layers includes two general fully connected layers (FC4, FC5) and a multi-domain fully connected layer (FC6), and the multi-domain fully connected layer is composed of a set of juxtaposed binary classification layers, and the binary classification scores of candidate pictures are obtained through the binary classification layers.

[0099] Step S05: Use the binary classification loss and the embedding loss to supervise the training of the binary classification prediction to obtain a trained convolutional-Transformer tracking network.

[0100]

[0101]

[0102] L = L cls + λL inst

[0103] Where: L cls represents the binary classification loss; L inst represents the embedding loss, which is used to overcome the situation where candidate samples are different from but overly similar to some classification categories; L represents the total loss function, and λ represents the weight (taking 0.1 in this example);

[0104] i = 1, 2, …, N, where N represents the total number of candidate samples; c ∈ {0, 1}, c = 1 indicates that the true label of the candidate sample is positive, and c = 0 indicates that the true label of the candidate sample is negative; y i represents the probability that the system correctly judges the true label of the i-th candidate sample; f i d represents the binary classification score of the i-th candidate sample on classification d, d = 1, 2, …, D, where D represents the total number of categories; [·] cd represents normalizing · for both c = 1 and c = 0 cases for classification d; [·] d represents normalizing · for c = 1 case for classification d; σ represents the softmax function.

[0105] Step S06: Input the video stream and the image to be tracked into the trained convolutional-Transformer tracking network, and track the image to be tracked in the video stream. The image to be tracked belongs to a certain frame picture or a part of a certain frame picture in the video stream.

[0106] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0107] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0108] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A RGB-T object tracking network method based on global attention convolutional-Transformer, characterized in that: It includes the following steps: (1) Obtain a training sample set from the RGB-T data stream. Each training sample includes a corresponding RGB image R and an infrared image T; (2) Construct a convolution-Transformer network to extract multimodal image features of RGB image R and infrared image T step by step, and aggregate the convolution features and Transformer features to obtain the RGB feature F vis and infrared signature F inf , including the following steps: (21) The convolutional branch network extracts features from the input image layer by layer through three convolutional blocks, obtaining the first, second, and third layer modal features f 1 , f 2 , f 3 respectively. Taking f 3 as the convolutional feature; the RGB image R passes through the convolutional branch network to obtain f 1 vis , the infrared image T passes through the convolutional branch network to obtain f 1 inf , (22) The Transformer branch network processes the input image to obtain X 0 , and then uses four groups of Transformer modules to extract features from the input image layer by layer, obtaining the first, second, third, and fourth layer global shared features X 1 , X 2 , X 3 , X 4 respectively. Finally, output processing is performed to obtain X, which is used as the Transformer feature; the RGB image R passes through the Transformer branch network to obtain and X vis , and the infrared image T passes through the Transformer branch network to obtain and X inf ; (23) Element-wise addition and X vis , and after layer normalization operation, RGB feature F is obtained vis ; Element-wise addition and X inf , and after layer normalization operation, infrared feature F is obtained inf ; (3) Construct a cross-attention fusion module to fuse the RGB feature F vis and the infrared feature F inf to strengthen the information interaction between the RGB image R and the infrared image T, and obtain the fused feature F out ; (4) Map the candidate samples to be predicted on the fused features, and perform binary classification prediction through a fully connected layer. The candidate samples to be predicted belong to a certain frame of picture or a part of a certain frame of picture in the RGB-T data stream; (5) Use binary classification loss and embedding loss to supervise the training of binary classification prediction, and obtain a trained convolutional-Transformer tracking network; (6) Input the video stream and the image to be tracked into the trained convolutional-Transformer tracking network, and track the image to be tracked in the video stream. The image to be tracked belongs to a certain frame of picture or a part of a certain frame of picture in the video stream.

2. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 1, characterized in that: In the step (21), the convolutional branch network consists of two convolutional branches with non-shared parameters. One convolutional branch is used to extract the features of the RGB image R, and the other convolutional branch is used to process the features of the infrared image T; Each convolutional branch consists of three sequentially connected convolutional blocks; The first convolutional block includes a filter, a ReLU activation function, a local response normalization layer and a max pooling layer, and is used to obtain the first layer of modal features. The filter is a filter with an output channel number of 96 and a convolutional kernel size of 7×7; The second convolutional block includes a filter, a ReLU activation function, a local response normalization layer and a max pooling layer, and is used to obtain the second layer of modal features. The filter is a filter with an output channel number of 256 and a convolutional kernel size of 5×5; The third convolutional block includes a filter and a ReLU activation function, and is used to obtain the third layer of modal features. The filter is a filter with an output channel number of 512 and a convolutional kernel size of 3×3.

3. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 1, characterized in that: In the step (22), the Transformer branch network obtains the global shared features according to the following method: (221) Perform input processing on the input image with an unfixed size, change the number of channels of the input image from 3 to 64, and obtain X 0 ; (222) Extract the global shared features through four stages. In each stage, a group of Transformer modules are used to process the input features; In the first, second, and third stages, first use the Transformer module group to process the input features, and then downsample the output of the Transformer module group. The downsampling result is used as the input feature of the next Transformer module group; In each Transformer module, the input features are first flattened, then passed through layer normalization operations to form the feature vector I. Then, through different linear transformation layers, linear mappings are performed to obtain the query set Q, the key set K, and the value set V. Next, self-attention operations are performed on the query set Q, the key set K, and the value set V. Then, the input features and the results of the self-attention operations are connected through residual connections. Finally, the output of the Transformer module is formed through layer normalization operations, which is expressed as: Q = L Q (I), K = L K (I), V = L V (I) Attention(Q,K,V)=softmamax(QK T )V SA(I) = Attention(Q, K, V) Where: L Q , L K , L V respectively represent the linear transformation layers of the query set, key set, and value set in the Transformer module; Attention represents the operation function of the self-attention mechanism, softmax represents the normalization exponential function for multi-classification, and SA represents the process function of self-attention in the Transformer module; (223) The outputs of the first, second, third, and fourth stages are X 1 , X 2 , X 3 , X 4 , and the output processing of X 4 results in X, that is, using the pooling layer to transform X 4 into X with the same size as f 3 to achieve alignment between features.

4. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 3, characterized in that: In the step (222), the extraction of global shared features is completed through four stages. In the first, second, third, and fourth stages, the number of Transformer modules in the Transformer module group is 2, 2, 2, and 1 respectively. The Transformer modules in the same Transformer module group are connected in sequence.

5. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 3, characterized in that: In the step (222), the extraction of global shared features is completed through four stages. In the first, second, and third stages, the input features are first processed using a Transformer module group, and then the output of the Transformer module group is downsampled. Through downsampling, the feature sizes of the outputs of the Transformer module groups in the first, second, and third stages are respectively adjusted to H, W, and C respectively represent the 0 length, width, and number of channels of X.

6. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 1, characterized in that: In the step (3), a cross-attention fusion module is constructed to strengthen the information interaction and fusion between the RGB image R and the infrared image T, including the following steps: (31) Obtain the RGB feature F through the linear transformation layer of the key set and value set vis of the key set and value set Obtain the infrared feature F through the linear transformation layer of the query set inf of the query set Perform cross-attention operation to obtain the cross feature CA(F vis ,F inf ) = Attention(Q inf ,K vis ,V vis ) Where: CA represents the operation function of cross-attention, and LN represents layer normalization operation; respectively represent calculating cross features the linear transformation layers of the query set, key set, and value set in the cross-attention branch; (32) Obtain the infrared feature F through the linear transformation layer of the key set and value set inf of the key set and value set Obtain the RGB feature F through the linear transformation layer of the query set vis of the query set Perform cross-attention operation to obtain cross features CA(F inf ,F vis ) = Attention(Q vis ,K inf ,V inf ) Wherein: respectively represent the linear transformation layers of the query set, key set, and value set in the cross-attention branch for calculating cross features ; (33) Perform a fusion operation to obtain the fused feature F out : where: Concat represents the operation function for feature fusion.

7. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 1, characterized in that: In the step (4), first, the candidate samples to be predicted are mapped to a set size on the fused feature F through the ROIAlign layer out , and then the binary classification scores of the candidate samples are predicted through a continuous set of fully connected layers.

8. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 7, characterized in that: The continuous set of fully connected layers includes two general fully connected layers and one multi-domain fully connected layer. The multi-domain fully connected layer is composed of a set of parallel binary classification layers. The binary classification scores of the candidate pictures are obtained through the binary classification layers.

9. The RGB-T object tracking network method based on global attention convolutional-Transformer according to claim 1, characterized in that: In the step (5), binary classification loss and embedding loss are used to supervise the training of binary classification predictions: L = L cls + λL inst Where: L cls represents the binary classification loss; L inst represents the embedding loss, which is used to overcome the situation where the candidate samples are different from but overly similar to certain classification categories; L represents the total loss function, and λ represents the weight; i = 1, 2, …, N, where N represents the total number of candidate samples; c ∈ {0, 1}, c = 1 indicates that the true label of the candidate sample is positive, and c = 0 indicates that the true label of the candidate sample is negative; y i represents the probability that the system correctly judges the true label of the i-th candidate sample; f i d represents the binary classification score of the i-th candidate sample on classification d, where d = 1, 2, …, D and D represents the total number of categories; [·] cd represents normalizing · for both cases of c = 1 and c = 0 for classification d; [·] d represents normalizing · for the case of c = 1 for classification d; σ represents the softmax function.

Citation Information

Patent Citations

  • Transform-based single target tracking method

    CN114266996A

  • Unsupervised RGB-T target tracking method based on attention multi-modal feature fusion

    CN114494354A