RGBT target tracking method and device based on feature reconstruction and cooperative interaction

By adopting the method of feature reconstruction and collaborative interaction in RGBT target tracking, and using the trained RGBT target tracking model to effectively integrate the dual-modal information, the problem of limited performance in the face of complex challenges is solved, and higher tracking accuracy and robustness are achieved.

CN120147363AActive Publication Date: 2025-06-13XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202510297758.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing RGBT target tracking methods have limited performance in the face of complex challenges such as scale changes and deformation, and when the target is blocked or thermal crossover occurs, there are significant differences in the information that can be provided by visible light and thermal infrared modes, resulting in insufficient robustness and adaptability.

Method used

The RGBT target tracking method based on feature reconstruction and collaborative interaction is adopted, and the target tracking is carried out through the trained RGBT target tracking model to effectively utilize and integrate RGBT bimodal information. The model includes feature extraction module, feature enhancement module, feature fusion module and tracking prediction module. Through technical means such as twin architecture, scale multi-directional joint module and hierarchical median fusion module, the feature information of the target on different scales is captured and feature fusion is performed to improve tracking accuracy and robustness.

Benefits of technology

It improves the accuracy and robustness of RGBT target tracking, can better cope with challenges in complex scenarios such as target scale changes and occlusion, and enhances the adaptability and practicality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147363A_ABST
    Figure CN120147363A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of visual target tracking, and discloses an RGBT target tracking method and device based on feature reconstruction and cooperative interaction. The RGBT target tracking method comprises the following steps: obtaining a visible light video sequence and a thermal infrared video sequence of a target to be tracked in a selected time period; taking the obtained visible light video sequence and the thermal infrared video sequence as input at the same time, and performing target tracking processing frame by frame by using a pre-trained RGBT target tracking model to obtain a target tracking result; the RGBT target tracking model comprises a feature extraction module, a feature enhancement module, a feature fusion module and a tracking prediction module. In the technical scheme disclosed by the invention, target tracking is performed by adopting the trained RGBT target tracking model, effective utilization and fusion of RGBT bimodal information are realized, and the accuracy and robustness of RGBT target tracking can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual object tracking, and particularly relates to an RGBT object tracking method and device based on feature reconstruction and collaborative interaction. Background Art

[0002] With the rapid development of artificial intelligence, visual object tracking, as one of the key directions in computer vision research, has received extensive attention in various fields in recent years and rich research results have been accumulated. Specifically, in the field of visual object tracking, by analyzing continuous frames or video sequences at different time points, information such as the motion trajectory, speed change, and behavior pattern of the target can be predicted. The acquisition of this information has extremely important practical significance for multiple key fields such as intelligent transportation, video surveillance, human-computer interaction, and national defense.

[0003] Currently, visual object tracking is still in the initial stage of the development of tracking algorithms. Most trackers rely on a single visible light modality, and the tracking performance needs to be improved when facing extreme environments. Given the complementary nature between different modality information, fusing multiple modality data can provide more comprehensive target features. With the continuous progress of sensor technology, multiple information sources have been proven to be able to effectively improve the tracking performance; in addition, the manufacturing cost of sensors has dropped significantly, and various sensors are increasingly widely used in daily life. Especially for thermal infrared sensors, the thermal infrared images generated by them are not affected by changes in illumination, and the target can be captured using thermal information, which makes the object tracking based on visible light and thermal infrared (RGB-Thermal, RGBT) one of the research hotspots.

[0004] In recent years, the exploration of RGBT object tracking algorithms has been continuously deepened, but there are still certain limitations; among them, when facing complex challenges such as scale changes and deformations, the effectiveness of existing methods is limited; in addition, when the target is occluded or there is a thermal crossover, there are significant differences in the information that the visible light and thermal infrared modalities can provide; in summary, how to effectively utilize the information of the two modalities to improve the robustness and adaptability of the RGBT object tracking technology has become an urgent technical problem to be solved. Summary of the Invention

[0005] The purpose of the present invention is to provide an RGBT object tracking method and device based on feature reconstruction and collaborative interaction to solve one or more of the above-mentioned technical problems. In the technical solution disclosed by the present invention, a trained RGBT object tracking model is used for object tracking, realizing the effective utilization and fusion of RGBT dual-modal information, and being able to improve the accuracy and robustness of RGBT object tracking.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] In the first aspect of the present invention, a method for RGBT object tracking based on feature reconstruction and collaborative interaction is provided, including the following steps:

[0008] Obtain the visible light video sequence and the thermal infrared video sequence of the selected time period of the target to be tracked;

[0009] Take the obtained visible light video sequence and thermal infrared video sequence as inputs at the same time, and use the pre-trained RGBT object tracking model to perform object tracking processing frame by frame to obtain the object tracking result;

[0010] Among them, the RGBT object tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module, and a tracking prediction module; among them,

[0011] The feature extraction module is a feature extraction network based on a Siamese architecture, which is used to input visible light modality data and thermal infrared modality data and perform feature extraction, and output initial feature maps of four branches; among them, the four branches are the visible light modality template branch, the visible light modality detection branch, the thermal infrared modality template branch, and the thermal infrared modality detection branch;

[0012] The feature enhancement module includes four-scale multi-directional joint modules, and each scale multi-directional joint module is used to input the initial feature map of one branch and perform feature enhancement, and output the enhanced feature map of the corresponding branch; among them, when the scale multi-directional joint module performs feature enhancement, it first performs multi-scale enhancement on the initial feature map, and then performs multi-directional joint interaction in the three directions of height, width, and channel to finally obtain the enhanced feature map;

[0013] The feature fusion module includes four-level median fusion modules, and each level median fusion module is used to input the enhanced feature map of one branch and fuse it layer by layer from the lowest level to the upper level to obtain the fused features of each level; add the fused features of each level element by element to obtain the fused feature map of the corresponding branch;

[0014] The tracking prediction module is used to input the fused feature maps of each branch and perform cross-correlation operations to output the object tracking result.

[0015] A further improvement of the present invention is that the specific steps for the scale multi-directional joint module to perform feature enhancement include:

[0016] Process the input initial feature map through parallel enhancement branches with different dilation rates to obtain multi-scale features;

[0017] The obtained multi-scale features are respectively subjected to average pooling and max pooling operations in the three directions of height, width, and channel. Then, the average pooling in the height direction is multiplied by the average pooling in the width direction, and the max pooling in the height direction is multiplied by the max pooling in the width direction to obtain the features processed in the three directions;

[0018] The features processed in the three directions are applied to a shared convolutional layer to obtain three attention weights;

[0019] The multi-scale features are weighted according to the three obtained attention weights, and the final enhanced feature map is obtained by element-wise addition.

[0020] A further improvement of the present invention lies in that in the step of obtaining multi-scale features by processing the input initial feature map through parallel enhancement branches with different dilation rates,

[0021] In the parallel enhancement branches with different dilation rates, the last branch uses global average pooling to extract global context features;

[0022] The multi-scale features are integrated features obtained by concatenating the results of each branch in the channel dimension.

[0023] A further improvement of the present invention lies in that in the step of applying the features processed in the three directions to a shared convolutional layer to obtain three attention weights,

[0024] The shared convolutional layer consists of two fully connected layers, a ReLU activation function, and a Sigmoid activation function;

[0025] The calculation expressions for the three attention weights are:

[0026] A 1 = MLP(HPool Avg (U c )·WPool Avg (U c ))

[0027] A 2 = MLP(HPool Max (U c )·WPool Max (U c ))

[0028] A 3 = MLP(CPool(U c ));

[0029] In the formula, A 1 , A 2 and A 3They are the attention weights in three directions respectively; MLP represents the shared convolutional layer; HPool is the pooling operation in the height direction; WPool is the pooling operation in the width direction; CPool is the pooling operation in the channel direction.

[0030] A further improvement of the present invention lies in that in the hierarchical median fusion module, in the step of fusing layer by layer from the lowest layer to the upper layer,

[0031] For each selected layer other than the lowest layer, first sum the original features of the selected layer and the fused features of the next layer of the selected layer element by element, then perform channel attention and spatial attention weighting processing through the median hybrid attention network, and finally sum the processing results of the median hybrid attention network and the original features of the selected layer element by element to obtain the fused features of the selected layer;

[0032] Among them, when the median hybrid attention network performs channel attention weighting processing, global max pooling, global average pooling, and global median pooling are adopted; the original features and the fused features of the lowest layer are kept consistent.

[0033] A further improvement of the present invention lies in that the fused features of the selected layer are expressed as:

[0034]

[0035] In the formula, is the fused feature of the selected layer; X n-1 is the feature of the selected layer; is the fused feature of the next layer of the selected layer; MHAN represents the median hybrid attention network; n is the total number of layers;

[0036] MHAN(X) = μ(η(Cat(ρ(X 输出 ), θ(X 输出 )), W));

[0037]

[0038] In the formula, μ represents the Sigmoid function; η represents the fully connected layer, W represents the corresponding weight information of the fully connected layer, ρ represents the max pooling in the channel dimension, θ represents the average pooling in the channel dimension; X 输出 is the output feature of the channel attention; X is the input feature; MLP(·) represents the multi-layer perceptron of the channel attention; ε represents the global max pooling in the spatial dimension, φ represents the global average pooling, represents the global median pooling;

[0039] The fused feature map of the corresponding branch is expressed as:

[0040]

[0041] A further improvement of the present invention lies in that

[0042] the tracking and prediction module includes a classification branch for predicting the category of each position in the response map and a prediction box regression branch for predicting the position of the target;

[0043] The total loss function during training is expressed as: L = L cls + λL loc ;

[0044] where λ is a hyperparameter; L cls is the classification loss, which is implemented by the cross-entropy function; L loc is the regression loss, which is used to predict the position of the target box and is represented by the L1 smooth loss of the standard coordinates;

[0045] In the tracking and prediction module, the maximum operation and the corresponding element summation operation are respectively performed on the classification branch and the prediction box regression branch to obtain the optimal classification response map and regression response map.

[0046] In the second aspect of the present invention, an RGBT target tracking device based on feature reconstruction and collaborative interaction is provided, including:

[0047] A data acquisition module for acquiring a visible light video sequence and a thermal infrared video sequence of the selected time period of the target to be tracked;

[0048] A target tracking module for using the acquired visible light video sequence and thermal infrared video sequence as inputs at the same time, and performing target tracking processing frame by frame using a pre-trained RGBT target tracking model to obtain a target tracking result;

[0049] wherein, the RGBT target tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module, and a tracking and prediction module; wherein,

[0050] The feature extraction module is a feature extraction network based on a siamese architecture, which is used to input visible light modality data and thermal infrared modality data and perform feature extraction, and output initial feature maps of four branches; wherein, the four branches are respectively a visible light modality template branch, a visible light modality detection branch, a thermal infrared modality template branch, and a thermal infrared modality detection branch;

[0051] The feature enhancement module includes four scale multi-directional joint modules, each of which is used to input the initial feature map of a branch, enhance the features, and output the enhanced feature map of the corresponding branch. When enhancing the features, the scale multi-directional joint module first performs multi-scale enhancement on the initial feature map, then performs multi-directional joint interaction in the three directions of height, width, and channel, and finally obtains the enhanced feature map.

[0052] The feature fusion module includes four hierarchical median fusion modules, each of which is used to input the enhanced feature map of a branch, fuse layer by layer from the lowest layer to the upper layer, and obtain the fused features of each layer. The fused features of each layer are added element by element to obtain the fused feature map of the corresponding branch.

[0053] The tracking and prediction module is used to input the fused feature maps of each branch, perform cross-correlation operations, and output the target tracking result.

[0054] In the third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the RGBT target tracking method based on feature reconstruction and collaborative interaction according to any one of the first aspects of the present invention.

[0055] In the fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the RGBT target tracking method based on feature reconstruction and collaborative interaction according to any one of the first aspects of the present invention.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] In the technical solution disclosed by the present invention, a trained RGBT object tracking model is used for object tracking. The RGBT object tracking model realizes the effective utilization and fusion of RGBT bimodal information through the collaborative work of four stages: feature extraction, feature enhancement, feature fusion, and tracking prediction, improves the accuracy and robustness of object tracking, and can better address the existing problems in challenges such as object scale changes and occlusions, providing strong technical support for practical applications. Further specifically and explanatorily, in the RGBT object tracking model disclosed by the present invention, the Scale Multi-directional Joint Module (SMJM) can effectively capture the spatial features of the object at different scales and synergistically enhance the feature expression ability through multi-directional attention, solving problems such as scale changes and deformations; the Level Median Fusion Module (LMFM) can fuse features of different levels layer by layer through the Median Hybrid Attention Network (MHAN), effectively addressing problems such as object occlusion and thermal crossover. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art; obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0059] Figure 1 is a schematic flowchart of an RGBT object tracking method based on feature reconstruction and collaborative interaction in an embodiment of the present invention;

[0060] Figure 2 is a schematic overall flowchart of an RGBT object tracking method based on feature reconstruction and collaborative interaction in a specific embodiment of the present invention;

[0061] Figure 3 is a schematic diagram of an RGBT object tracking model in a specific embodiment of the present invention;

[0062] Figure 4 is a schematic diagram of the Scale Multi-directional Joint Module in a specific embodiment of the present invention;

[0063] Figure 5 is a schematic diagram of the Level Median Fusion Module in a specific embodiment of the present invention;

[0064] Figure 6 is a schematic diagram of the Median Hybrid Attention Network in a specific embodiment of the present invention;

[0065] Figure 7 It is a schematic diagram for evaluating the accuracy performance of the tracking method according to the embodiments of the present invention on the GTOT test set;

[0066] Figure 8 It is a schematic diagram for evaluating the success rate performance of the tracking method according to the embodiments of the present invention on the GTOT test set;

[0067] Figure 9 It is a schematic diagram for evaluating the accuracy performance of the tracking method according to the embodiments of the present invention on the RGBT234 test set;

[0068] Figure 10 It is a schematic diagram for evaluating the success rate performance of the tracking method according to the embodiments of the present invention on the RGBT234 test set;

[0069] Figure 11 It is a schematic diagram of an RGBT object tracking device based on feature reconstruction and collaborative interaction in the embodiments of the present invention. Detailed implementation manners

[0070] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments are only a part rather than all of the embodiments of the present invention.

[0071] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0072] Please refer to Figure 1 , an RGBT object tracking method based on feature reconstruction and collaborative interaction provided by the embodiments of the present invention includes the following steps:

[0073] Step 1, obtaining a visible light video sequence and a thermal infrared video sequence of a selected time period of the target to be tracked;

[0074] Step 2, taking the visible light video sequence and the thermal infrared video sequence obtained in Step 1 as inputs at the same time, and performing object tracking processing frame by frame using a pre-trained RGBT object tracking model to obtain an object tracking result;

[0075] Among them, the RGBT target tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module, and a tracking prediction module;

[0076] The feature extraction module is a feature extraction network based on a siamese architecture, which is used to input visible light modality data and thermal infrared modality data, perform feature extraction, and output initial feature maps of four branches; among them, the four branches are a visible light modality template branch, a visible light modality detection branch, a thermal infrared modality template branch, and a thermal infrared modality detection branch;

[0077] The feature enhancement module includes four-scale multi-directional joint modules, and each scale multi-directional joint module is respectively used to input the initial feature map of one branch, perform feature enhancement, and output the enhanced feature map of the corresponding branch; among them, when the scale multi-directional joint module performs feature enhancement, it first performs multi-scale enhancement on the initial feature map, and then performs multi-directional joint interaction in the three directions of height, width, and channel to finally obtain the enhanced feature map;

[0078] The feature fusion module includes four-level median fusion modules, and each level median fusion module is respectively used to input the enhanced feature map of one branch and fuse it layer by layer from the lowest level to the upper level to obtain the fused features of each level; add the fused features of each level element by element to obtain the fused feature map of the corresponding branch;

[0079] The tracking prediction module is used to input the fused feature maps of each branch, perform cross-correlation operations, and output the target tracking result.

[0080] In the technical solution provided by the embodiment of the present invention, the RGBT target tracking model uses dual-modal information of RGB (red, green, and blue three-channel color image) and T (thermal infrared image) for target tracking. Through the collaborative work of four stages: feature extraction, feature enhancement, feature fusion, and tracking prediction, this model realizes the effective utilization and fusion of dual-modal information. This fusion strategy not only improves the accuracy of target tracking but also enhances the robustness of the model, enabling it to better cope with challenges in complex scenarios such as target scale changes and occlusions.

[0081] Specifically and explanatorily, the main function of the scale multi-directional joint module (SMJM) is to capture the spatial features of the target at different scales, which can effectively handle the scale change problem of the target caused by distance changes or perspective changes during the tracking process. By enhancing the feature expression ability, it improves the robustness of the model to target deformation.

[0082] Specifically and explanatorily, the main function of the hierarchical median fusion module (LMFM) is to fuse features of different levels layer by layer. Through the median hybrid attention network (MHAN), it combines the stability of median filtering and the flexibility of the attention mechanism to fuse features layer by layer, which can effectively address the problem of target occlusion. Because features of different levels may contain target information before and after occlusion, fusing this information helps to restore the target features of the occluded part.

[0083] In a specific embodiment of the present invention, the specific steps for the scale multi-directional joint module to perform feature enhancement include:

[0084] First, the input initial feature map passes through parallel enhancement branches with different dilation rates to obtain multi-scale features that can better adapt to targets of different sizes;

[0085] Then, the multi-scale features are respectively subjected to average pooling and max pooling operations in the three directions of height, width, and channel. Multiply the average pooling in the height direction by the average pooling in the width direction to highlight the useful information in the input features and reduce redundancy; similarly, multiply the max pooling in the height direction by the max pooling in the width direction, and finally obtain the features processed in the three directions;

[0086] After that, apply a shared convolutional layer to the features processed in the three directions to obtain three attention weights;

[0087] Finally, weight the multi-scale features according to the attention weights and obtain the final output feature map by element-wise addition, so that the tracking model can better adapt to target scale changes and more comprehensively express the useful information of the target.

[0088] In an embodiment of the present invention, in the step of the hierarchical median fusion module performing feature fusion, the process of layer-by-layer fusion is as follows: First, starting from the lowest layer, add the features of this layer to the features of the previous layer; then input the integrated features into the median hybrid attention network for channel attention and spatial attention weighting processing, enabling the model to adaptively adjust the weights and more accurately extract and utilize useful information; different from the usual channel attention and spatial attention, the present invention further introduces global median pooling on the basis of global max pooling and global average pooling to suppress and reduce noise, because when there is obvious noise in the input features, these noises may affect the accuracy of the fusion result; then add the initial input features of this layer to obtain the fused features of this layer; finally, perform element-wise addition on the fused features obtained for each layer to obtain the final output feature map.

[0089] Please refer to Figure 2 and Figure 3, in a specific embodiment of the present invention, a RGBT object tracking method based on feature reconstruction and collaborative interaction is provided, which specifically includes the following steps:

[0090] Step S0: Construct a dataset, including: using the visible light dataset GOT10K as the pre-training dataset, the RGBT tracking dataset LasHeR as the training set, and GTOT and RGBT234 as the test sets; preprocess the LasHeR dataset before training. In a specific exemplary technical solution, each frame of the image is cropped into a pixel block of 511×511 size, and data augmentation techniques (including flipping, rotation, scaling, etc.) are used on it to improve the robustness and generalization ability of the model.

[0091] Step S1: Construct an object tracking model, including: the model includes a feature extraction network based on a siamese architecture, a feature enhancement module, a feature fusion module, and a tracking prediction module, as Figure 3 shown.

[0092] Step S2: Pre-training and fine-tuning of the model, including: using the pre-training dataset (GOT10K) to pre-train the feature extraction network, training it with the stochastic gradient descent method (SGD) until the loss value converges to obtain a pre-trained model; then using the training dataset (LasHeR) to fine-tune the pre-trained model, training it with SGD until the loss value converges to obtain a trained model;

[0093] Step S3: The stage of obtaining the tracking result, including: obtaining the first frame of the visible light and thermal infrared image sequences in the RGBT test dataset as the template image, and after selecting the object to be tracked, inputting it into the trained RGBT object tracking model based on feature reconstruction and collaborative interaction to generate the object tracking result.

[0094] Please refer to Figures 4 to 6 , in the embodiment of the present invention, when constructing the feature extraction network, the network adopts a multimodal symmetric form and is composed of 4 deep convolutional networks ResNet50, that is, two groups of siamese networks; the results of the last three layers of ResNet50 are used as the output results, and then the output results are transmitted to the downsampling block to adjust the size and number of channels of the feature map. When constructing the feature enhancement module, the embodiment of the present invention designs a scale multi-directional joint module, as Figure 4 shown. When constructing the feature fusion module: the embodiment of the present invention designs a hierarchical median fusion module, as Figure 5 shown; it includes a median hybrid attention network, as Figure 6 shown.

[0095] In an exemplary alternative technical solution of an embodiment of the present invention, when constructing a tracking and prediction network, a tracking result response map is obtained by performing a cross-correlation operation, which includes 4 outputs: classification score response maps of the predicted target categories of the two modalities respectively, and a predicted box regression response map for predicting the target position. For the classification branch, the best scoring classification is determined by comparing and selecting the maximum classification scores in the two modalities; for the regression branch, the corresponding elements of the predicted box regression response maps of the two modalities are summed to obtain a comprehensively predicted target position.

[0096] Please refer to Figure 2 , in step S2 of the above embodiment, the pre-training and fine-tuning of the model include the following steps:

[0097] Step S2-1: The feature extraction network includes 4 ResNet50s with the same structure. Every 2 ResNet50s form a group of the feature extraction network (representing one modality). The feature extraction network of a set of siamese architectures is pre-trained using a visible light dataset (GOT10K), and the model is trained using an SGD optimizer until the model loss value converges. Finally, the pre-trained model is saved;

[0098] Step S2-2: Based on the pre-trained model parameters saved in step S2-1, first, the model parameters are copied and extended into two groups to adapt to the feature extraction network required for the dual modality. Then, the processed pre-trained parameters are used as input to initialize the feature extraction network. The first two layers of the ResNet50 are frozen, and the overall model is fine-tuned using an RGBT training dataset (LasHeR). At the same time, the learning rate of the network is reduced, and the model is trained using an SGD optimizer until the model loss value converges. Finally, the trained model is obtained and saved.

[0099] Please see Figure 2 , in step S3 of the above embodiment, obtaining the templates of the targets to be tracked in the visible light and thermal infrared images of the RGBT test datasets (GTOT, RGBT234) and performing tracking includes the following steps:

[0100] Step S3-1: The template image is centered on the target selected from the first frame of the visible light and thermal infrared image sequences, and cropped to a size of 127×127 pixels. The subsequent frames are search images, and the search images are of size 255×255 pixels;

[0101] Step S3-2: The template images and search images of visible light and infrared obtained in step S3-1 are respectively input into the feature extraction network trained in step S2-2 to obtain multi-layer template features and search features of visible light and thermal infrared;

[0102] Step S3-3: Use the scale multi-directional joint module to perform scale feature enhancement on the multi-layer visible light and thermal infrared template features and visible light and infrared search features obtained in step S3-2, and generate multi-layer enhanced visible light and thermal infrared template features and search features.

[0103] Step S3-4: Use the hierarchical median fusion module to perform homologous fusion on the multi-layer enhanced visible light and thermal infrared template features and search features in each branch obtained in step S3-3, complete the interaction between the feature information of each modal level, and generate template features and search features after fusion of different modalities.

[0104] Step S3-5: Perform cross-correlation operations pairwise on the template features and search features after fusion of each modality to obtain a multi-layer response map for tracking prediction;

[0105] Step S3-6: Perform maximum value operation and element-wise summation operations on the classification branch and the regression branch respectively, aiming to obtain the optimal classification response map and regression response map. This processing method can extract the best tracking information from the classification accuracy and the accuracy of position prediction respectively.

[0106] As Figure 4 shown, the scale multi-directional joint module provided by the embodiment of the present invention aims to use dilated convolution and multi-directional attention to cooperate to capture target information at different scales and enhance the feature expression ability of the target.

[0107] First, extract features of different scales from the input features through five parallel branches. Each branch is configured with a different dilation rate, and the last branch uses global average pooling to extract global context features, which helps to enhance the model's perception ability of the overall image information. Then, concatenate the results of each branch in the channel dimension to obtain an integrated feature. Assuming the input feature is The obtained integrated feature U c is as follows:

[0108]

[0109] Among them, and respectively represent convolution operations with different dilation rates, and W 1 , W 2 , W 3 and W 4 represent the corresponding weight information of the convolution.

[0110] Next, average pooling and max pooling operations are performed on the integrated features in the height direction, width direction, and channel direction respectively. Then, the average pooling in the height direction is multiplied by the average pooling in the width direction to highlight the useful information in the input features and reduce redundancy. Similarly, the max pooling in the height direction is multiplied by the max pooling in the width direction.

[0111] Then, the feature maps of the three branches are processed by a shared convolutional layer to obtain three attention weights. The shared convolutional layer consists of two fully connected layers, a ReLU activation function, and a Sigmoid activation function, and is used to reduce and restore the channel dimension.

[0112] A 1 = MLP(HPool Avg (U c )·WPool Avg (U c ))

[0113] A 2 = MLP(HPool Max (U c )·WPool Max (U c ))

[0114] A 3 = MLP(CPool(U c ))

[0115] Among them, MLP represents the shared convolutional layer, HPool is the pooling operation in the height direction, WPool is the pooling operation in the width direction, CPool is the pooling operation in the channel direction, and A 1 , A 2 and A 3 are the attention weights in the three directions respectively.

[0116] Finally, the input multi-scale integrated features are weighted according to the attention weights, and the final output feature map is obtained by element-wise addition.

[0117] As Figure 5 shown, the hierarchical median fusion module provided by the embodiment of the present invention gradually integrates the feature information of different levels through a series of feature refinement units, and gradually fuses the low-level features with the high-level features.

[0118] The specific operation is as follows: First, the features of the next level are added to the upper-level features X n-1 , then processed by a median hybrid attention network including multiple refinement units, and finally the original features X n-1 of this layer are added to obtain the fused features That is:

[0119]

[0120] Note that during the fusion process, the original features and the fused features of the bottommost branch remain the same.

[0121] Among them, the framework of the median hybrid attention network is as Figure 6 shown, which is different from the usual channel attention and spatial attention. Based on global max pooling and global average pooling, we further introduce global median pooling to suppress and reduce noise. Because when there is obvious noise in the input features, these noises may affect the accuracy of the fusion result. The specific formula is:

[0122]

[0123] MHAN(X) = μ(η(Cat(ρ(X 1 ), θ(X 1 ))), W))

[0124] Among them, X is the input feature, X 1 is the output feature of channel attention, MLP(·) represents the multi-layer perceptron of channel attention, ε represents global max pooling in the spatial dimension, φ represents global average pooling, represents global median pooling, ρ represents max pooling in the channel dimension, θ represents average pooling in the channel dimension, η represents the fully connected layer, W represents the corresponding weight information of the fully connected layer, and μ represents the Sigmoid function.

[0125] Finally, we add the fused features of each layer element-wise to form a final fused feature

[0126]

[0127] In the exemplary solution of the embodiment of the present invention, the tracking prediction module includes a classification branch for predicting the category of each position in the prediction response map and a prediction box regression branch for predicting the target position, and each branch participates in the calculation of the loss; among them, the total loss function of the network can be expressed as:

[0128] L = L cls + λL loc ;

[0129] Among them, λ is a hyperparameter; L cls is the classification loss, which is implemented by the cross-entropy function; L loc is the regression loss, which is used to predict the position of the target box and is represented by the L1 smooth loss of the standard coordinates.

[0130] In step S3-6, the processed classification response map uses a cosine window and a penalty mechanism to rearrange the scores of the prediction boxes to find the best prediction result; among them, the cosine window suppresses large displacements, and the penalty mechanism suppresses large changes in size and ratio. The specific expression of the penalty is:

[0131]

[0132] where k is a hyperparameter, r is the aspect ratio, r′ is the aspect ratio of the last frame, s is the target size of the current frame, and s′ is the target size of the last frame.

[0133] The final classification score map cls′ is expressed as:

[0134] cls′ = (1 - α) * cls * penalty + α * H;

[0135] where α is a hyperparameter, cls is the initial prediction score map, and H is the cosine window.

[0136] Exemplary hyperparameter settings are: trained using the SGD optimizer, with a momentum parameter of 0.9, a weight decay of 0.0005, the number of training epochs set to 20, and the initial learning rate set to 10 -2 , and the final learning rate set to 10 -5 , with exponential decay between each training epoch. These hyperparameter settings are designed to ensure that the model can converge efficiently and stably.

[0137] In summary, in the tracking model of the embodiment of the present invention (each modality includes two branches, namely a template branch and a detection branch, so there are a total of four input branches. The template branch is used to store the initial feature information of the target object as a benchmark for matching; the detection branch is used to search for the target area in the current frame that is most similar to the template feature to complete target localization), the input data first passes through the feature extraction part to obtain the feature maps of each input branch (each branch contains three levels of features, and the combination of multiple levels can achieve global and local complementarity and improve the comprehensive performance of the model); then these feature maps are used as inputs and loaded into the scale multi-directional joint module to enhance the features and obtain enhanced feature maps so that the features can adapt to changes in different scales; then the enhanced feature maps are input into the hierarchical median fusion module to fuse the features of multiple levels to obtain the fused feature maps; finally, cross-correlation operations are performed on the fused features obtained from different branches of each modality to obtain the tracking result. The tracking result is the border data of the prediction box; specifically, the accuracy and success rate error curve graphs of the prediction box and the ground truth box are usually used as the tracking result to reflect the tracking performance of the model, such as Figure 7 、 8, 9, and 10 show the accuracy and success rate performance evaluation results of the GTOT and RGBT234 test sets.

[0138] In summary, the embodiments of the present invention disclose an RGBT object tracking method based on feature reconstruction and collaborative interaction, aiming to further improve the tracking performance by utilizing the complementary characteristics of multi-source information. For problems such as scale changes and deformations, the present invention adopts a scale multi-directional joint module to capture the feature information of the target at different scales, including the detailed information and extensive context information in the image, and uses multi-directional attention collaboration to further enhance the feature expression ability. For problems such as target occlusion and thermal crossover, the present invention adopts a hierarchical median fusion module to fuse the feature information of different modalities and different levels in each branch through a median hybrid attention network, realizing the collaboration between the texture information represented by the shallow features and the semantic information represented by the deep features to better address the above problems. Comparative experiments are conducted with other mainstream RGBT object tracking algorithms on two commonly used datasets GTOT and RGBT234 in the field of RGBT object tracking to illustrate the superiority of the present invention.

[0139] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.

[0140] Please refer to Figure 11 , in the embodiments of the present invention, the embodiments of the present invention provide an RGBT object tracking device based on feature reconstruction and collaborative interaction, including:

[0141] A data acquisition module, configured to acquire a visible light video sequence and a thermal infrared video sequence of a selected time period of the target to be tracked;

[0142] An object tracking module, configured to use the acquired visible light video sequence and thermal infrared video sequence as inputs at the same time, and perform object tracking processing frame by frame using a pre-trained RGBT object tracking model to obtain an object tracking result;

[0143] Wherein, the RGBT object tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module, and a tracking prediction module; wherein,

[0144] The feature extraction module is a feature extraction network based on a Siamese architecture, configured to input visible light modality data and thermal infrared modality data and perform feature extraction, and output initial feature maps of four branches; wherein, the four branches are a visible light modality template branch, a visible light modality detection branch, a thermal infrared modality template branch, and a thermal infrared modality detection branch;

[0145] The feature enhancement module includes four scale multi-directional joint modules, and each scale multi-directional joint module is respectively used to input the initial feature map of a branch and perform feature enhancement, and output the enhanced feature map of the corresponding branch. Among them, when the scale multi-directional joint module performs feature enhancement, it first performs multi-scale enhancement on the initial feature map, and then performs multi-directional joint interaction in the three directions of height, width, and channel to finally obtain the enhanced feature map.

[0146] The feature fusion module includes four hierarchical median fusion modules, and each hierarchical median fusion module is respectively used to input the enhanced feature map of a branch and fuse it layer by layer from the lowest level to the upper level to obtain the fused features of each level; the fused features of each level are added element by element to obtain the fused feature map of the corresponding branch.

[0147] The tracking and prediction module is used to input the fused feature maps of each branch and perform cross-correlation operations to output the target tracking result.

[0148] In an embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to execute the operations of the RGBT target tracking method based on feature reconstruction and collaborative interaction.

[0149] In an embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM (Random Access Memory) or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the RGBT target tracking method based on feature reconstruction and collaborative interaction in the above embodiments.

[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, optical memories, etc.) containing computer-usable program codes.

[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the process in Figure 1one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A RGBT target tracking method based on feature reconstruction and collaborative interaction, characterized in that: The following steps are involved: Obtaining a visible light video sequence and a thermal infrared video sequence of a selected time period of the target to be tracked; The acquired visible light video sequence and thermal infrared video sequence are used as input at the same time, and the pre-trained RGBT target tracking model is used to perform target tracking processing frame by frame to obtain the target tracking result; The RGBT target tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module and a tracking prediction module; The feature extraction module is a feature extraction network based on a twin architecture, which is used to input visible light modal data and thermal infrared modal data and perform feature extraction, and output an initial feature map of four branches; wherein the four branches are a visible light modal template branch, a visible light modal detection branch, a thermal infrared modal template branch, and a thermal infrared modal detection branch; The feature enhancement module includes four scale multi-directional joint modules, each of which is used to input an initial feature map of a branch and perform feature enhancement, and output an enhanced feature map of the corresponding branch; wherein, when the scale multi-directional joint module performs feature enhancement, the initial feature map is firstly multi-scale enhanced, and then multi-directional joint interaction is performed in three directions of height, width and channel, and finally the enhanced feature map is obtained; The feature fusion module includes four hierarchical median fusion modules, each hierarchical median fusion module is used to input the enhanced feature map of a branch and fuse it from the lowest level to the upper level layer by layer to obtain the fused features of each level; the fused features of each level are element-wise added to obtain the fused feature map of the corresponding branch; The tracking prediction module is used to input the fused feature maps of each branch and perform cross-correlation operations, and output the target tracking result.

2. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 1, characterized in that: The specific steps of the scale multi-directional joint module for feature enhancement include: The input initial feature map is processed by parallel enhancement branches with different void ratios to obtain multi-scale features; The obtained multi-scale features are subjected to average pooling and maximum pooling operations in the three directions of height, width and channel, and then the average pooling in the height direction is multiplied by the average pooling in the width direction, and the maximum pooling in the height direction is multiplied by the maximum pooling in the width direction to obtain the features processed in the three directions; Apply the shared convolution layer to the features processed in the three directions to obtain three attention weights; The multi-scale features are weighted according to the three attention weights obtained, and the final enhanced feature map is obtained by adding the corresponding elements.

3. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 2 is characterized in that: In the step of processing the input initial feature map through parallel enhancement branches with different void rates to obtain multi-scale features, the last branch of the parallel enhancement branches with different void rates uses global average pooling to extract global context features; The multi-scale feature is an integrated feature obtained by concatenating the results of each branch in the channel dimension.

4. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 2, characterized in that: In the step of applying the features processed in three directions to a shared convolutional layer to obtain three attention weights, The shared convolutional layer consists of two fully connected layers, a ReLU activation function and a Sigmoid activation function; The calculation expressions of the three attention weights are: A1=MLP(HPool Avg (U c )·WPool Avg (U c )) <h2 style=";text-align:left;direction:ltr">A2=MLP(HPool<h2 style=";text-align:left;direction:ltr"> Max <h2 style=";text-align:left;direction:ltr"> (U<h2 style=";text-align:left;direction:ltr"> c <h2 style=";text-align:left;direction:ltr"> )·WPool<h2 style=";text-align:left;direction:ltr"> Max <h2 style=";text-align:left;direction:ltr"> (U<h2 style=";text-align:left;direction:ltr"> c <h2 style=";text-align:left;direction:ltr"> )) A3=MLP(CPool(U c )); In the formula, A1, A2 and A3 are the attention weights in three directions respectively; MLP stands for shared convolutional layer; HPool is a pooling operation in the height direction; WPool is a pooling operation in the width direction; CPool is a pooling operation in the channel direction.

5. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 1, characterized in that: In the hierarchical median fusion module, in the step of fusing from the lowest level to the upper level layer by layer, For each selected level other than the lowest level, the original features of the selected level are first added element by element with the fused features of the next level of the selected level, and then the median mixed attention network is used to perform channel attention and spatial attention weighted processing. Finally, the median mixed attention network processing result is added element by element with the original features of the selected level to obtain the fused features of the selected level. Among them, when the median hybrid attention network performs channel attention weighted processing, global maximum pooling, global average pooling and global median pooling are adopted; the original features of the lowest level are kept consistent with the fused features.

6. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 5, characterized in that: The fused features of the selected level are expressed as: In the formula, is the fused features of the selected level; X n-1 is the feature of the selected level; is the fused features of the next level of the selected level; MHAN stands for median hybrid attention network; n is the total number of levels; MHAN(X)=μ(η(Cat(ρ(X 输出 ),θ(X 输出 )),W)); In the formula, μ represents the Sigmoid function; η represents the fully connected layer, W represents the corresponding weight information of the fully connected layer, ρ represents the maximum pooling in the channel dimension, and θ represents the average pooling in the channel dimension; X 输出 is the output feature of channel attention; X is the input feature; MLP(·) represents the multi-layer perceptron of channel attention; ε represents the global maximum pooling in the spatial dimension, φ represents the global average pooling, represents global median pooling; The fused feature map of the corresponding branch It is expressed as:

7. The RGBT target tracking method based on feature reconstruction and collaborative interaction according to claim 1, characterized in that: The tracking prediction module includes a classification branch for predicting each position category in the response map and a prediction box regression branch for predicting the target position; The total loss function during training is expressed as: L = L cls +λL loc ; Where λ is a hyperparameter; L cls is the classification loss, which is implemented by the cross entropy function; L loc It is the regression loss used to predict the position of the target box and is represented by the L1 smoothing loss of the standard coordinates; In the tracking prediction module, the maximum value operation and the corresponding element summation operation are respectively performed on the classification branch and the prediction frame regression branch to obtain the optimal classification response map and regression response map.

8. An RGBT target tracking device based on feature reconstruction and collaborative interaction, characterized in that: include: A data acquisition module, used to acquire a visible light video sequence and a thermal infrared video sequence of a selected time period of a target to be tracked; The target tracking module is used to take the acquired visible light video sequence and thermal infrared video sequence as input at the same time, and use the pre-trained RGBT target tracking model to perform target tracking processing frame by frame to obtain the target tracking result; The RGBT target tracking model includes: a feature extraction module, a feature enhancement module, a feature fusion module and a tracking prediction module; The feature extraction module is a feature extraction network based on a twin architecture, which is used to input visible light modal data and thermal infrared modal data and perform feature extraction, and output an initial feature map of four branches; wherein the four branches are a visible light modal template branch, a visible light modal detection branch, a thermal infrared modal template branch, and a thermal infrared modal detection branch; The feature enhancement module includes four scale multi-directional joint modules, each of which is used to input an initial feature map of a branch and perform feature enhancement, and output an enhanced feature map of the corresponding branch; wherein, when the scale multi-directional joint module performs feature enhancement, the initial feature map is firstly multi-scale enhanced, and then multi-directional joint interaction is performed in three directions of height, width and channel, and finally the enhanced feature map is obtained; The feature fusion module includes four hierarchical median fusion modules, each hierarchical median fusion module is used to input the enhanced feature map of a branch and fuse it from the lowest level to the upper level layer by layer to obtain the fused features of each level; the fused features of each level are element-wise added to obtain the fused feature map of the corresponding branch; The tracking prediction module is used to input the fused feature maps of each branch and perform cross-correlation operations, and output the target tracking result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the RGBT target tracking method based on feature reconstruction and collaborative interaction according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the RGBT target tracking method based on feature reconstruction and collaborative interaction according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Text generation image algorithm based on multi-scale features

    CN115862039A

  • RGBT target tracking method based on cross-modal feature self-enhancement and step-by-step fusion

    CN118037769A

  • Visual target tracking method and apparatus based on deeply and densely connected neural network

    US20200327679A1

Cited By

  • Cross-modal target tracking method based on collaborative strategy and related device

    CN121120693A