A cross-modal target tracking method based on a cooperative strategy and related devices
Patent Information
- Application Number
- CN202511192432.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2045-08-25
AI Technical Summary
然而,上述现有跨模态目标跟踪方法在特征建模和特征融合上仍面临一定挑战;其中,一方面,许多方法依赖单尺度或粗粒度特征进行匹配,缺乏对多尺度上下文的感知能力,难以有效应对目标的快速缩放与形变;另一方面,部分深层网络融合方案在整合多源信息时易引入特征冗余或语义偏移,导致融合特征的判别性能下降,从而限制了整体跟踪效果的提升
[0042]This invention discloses a cross-modal target tracking method based on a collaborative strategy. This method introduces a multi-scale channel adaptive enhancement module (MCAEM) and a hierarchical progressive attention fusion module (HPAFM) to achieve efficient collaborative fusion of visible light and thermal infrared modal features. This effectively addresses challenges such as scale changes, deformation, and texture blurring, improving tracking stability and robustness under occlusion, modal interference, and low-light conditions. Specifically, addressing the problem of insufficient modal information fusion in existing methods, this invention constructs a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module. By introducing multi-scale structure perception and dynamic attention mechanisms, efficient collaborative fusion between infrared and visible light features is achieved. This design enhances the complementary expressive ability between modalities and strengthens the discriminative power of the fused features, thereby significantly improving tracking accuracy and robustness. To elaborate further, the multi-scale channel adaptive enhancement module can capture the semantic and spatial features of the target at different scales through collaborative modeling with channel attention mechanism and multi-branch dilated convolution, thereby enhancing the perception of key channels and scale regions and effectively addressing challenges such as scale changes, deformation, and texture blurring. In addition, the hierarchical progressive attention fusion module integrates shallow and deep features layer by layer with the help of convolutional attention mechanism, highlights the target region through explicit modeling of channels and space, strengthens multi-layer semantic consistency, and improves the tracking stability and robustness of the network under occlusion, modal interference, and low-light conditions.
Smart Images

Figure CN121120693B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and image processing technology, and relates to multimodal perception and target tracking technology, particularly to a cross-modal target tracking method and related apparatus based on a cooperative strategy. Background Technology
[0002] In recent years, target tracking technology has been widely used in fields such as intelligent security, autonomous driving, and aerospace remote sensing. Its core objective is to locate and track a specified target in real time in continuous frame images. Among them, tracking methods based on Siam networks, such as SiamFC and SiamRPN, have become one of the mainstream solutions. These methods achieve good tracking performance while maintaining a certain level of real-time performance by constructing a template graph and calculating the feature similarity between the template graph and the search graph.
[0003] To further enhance adaptability in extreme environments, the academic community has proposed cross-modal target tracking schemes; among them, visible light and thermal infrared (RGB-Thermal, TIR) fusion tracking algorithms have received widespread attention in recent years. Specifically, existing RGBT tracking methods can be broadly categorized into two types: one directly stitches or weights together visible light (Red, Green, Blue, RGB) and thermal infrared (TIR) image features; the other employs attention mechanisms or modal adaptive modules to improve the fusion effect. However, these existing cross-modal target tracking methods still face certain challenges in feature modeling and feature fusion. On the one hand, many methods rely on single-scale or coarse-grained features for matching, lacking the ability to perceive multi-scale context and struggling to effectively cope with rapid scaling and deformation of targets. On the other hand, some deep network fusion schemes easily introduce feature redundancy or semantic shifts when integrating multi-source information, leading to a decrease in the discriminative performance of the fused features, thus limiting the improvement of overall tracking performance.
[0004] In summary, although cross-modal tracking technology has certain advantages in terms of environmental adaptability, there is still considerable room for improvement in its fusion efficiency and feature modeling capabilities, especially in complex scenarios (such as occlusion and scale transformation scenarios). Therefore, it is necessary to propose a fusion strategy with better structural adaptability and information collaboration capabilities to fully unleash the potential of multimodal perception. Summary of the Invention
[0005] The purpose of this invention is to provide a cross-modal target tracking method and related apparatus based on a cooperative strategy to solve one or more of the aforementioned technical problems. The technical solution disclosed in this invention can effectively address challenges such as scale changes, deformation, and texture blurring, improving tracking stability and robustness under occlusion, modal interference, and low-light conditions.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a cross-modal target tracking method based on a cooperative strategy, comprising the following steps:
[0008] Get the template image and the search image in the current frame;
[0009] Based on the acquired template image and the current frame search image, cross-modal target tracking is performed using the trained cross-modal target tracking model to obtain cross-modal target tracking results;
[0010] The cross-modal target tracking model includes four branch networks and a multimodal region proposal network. The four branch networks are respectively used as inputs for a visible light template image, a thermal infrared template image, a visible light search image, and a thermal infrared search image, and respectively output visible light template enhancement features, thermal infrared template enhancement features, visible light search enhancement features, and thermal infrared search enhancement features. The multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates target candidate regions through a fully convolutional network, which serve as the cross-modal target tracking result.
[0011] The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. Each extracted layer feature is then processed by the multi-scale channel adaptive enhancement module through channel attention enhancement and parallel dilated convolution to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are then fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and to obtain the final global enhanced features containing multi-level collaborative semantics.
[0012] A further improvement to the technical solution of this invention is that the multimodal region proposal network includes: a visible light branch and a thermal infrared branch; wherein,
[0013] Both the visible light branch and the thermal infrared branch include classification branches and regression branches;
[0014] In the classification and regression branches, the input enhancement features are adjusted through convolution operations, and the adjusted template features are used as convolution kernels to perform cross-correlation operations with the adjusted search features to generate visible light classification branch response maps, visible light regression branch response maps, thermal infrared classification branch response maps, and thermal infrared regression branch response maps. The visible light and thermal infrared corresponding branch response maps are weighted and fused to obtain the final classification response map and regression response map, and the cross-modal target tracking results are obtained based on the final response maps.
[0015] A further improvement to the technical solution of this invention lies in the fact that, in the multi-scale channel adaptive enhancement module, the steps of performing channel attention enhancement and parallel dilated convolution processing to obtain enhanced features that fuse features of different scales specifically include:
[0016] First, the input selected layer features are processed by global max pooling and global average pooling to extract channel-level global context descriptions. The results of the two pooling processes, global max pooling and global average pooling, are then concatenated to obtain the concatenated result.
[0017] Then, the splicing result is input into a two-layer fully connected network. The first fully connected network uses ReLU activation to capture non-linear channel dependencies, and the second fully connected network uses Sigmoid activation to generate normalized channel attention weights. The channel attention weights are multiplied with the input selected layer features channel by channel to obtain weighted features.
[0018] Finally, based on the weighted features, multi-path parallel dilated convolution branches are used to extract multi-scale information. The outputs of each branch are added element-wise along the channel dimension and fused into a dynamic perceptual feature map with multi-scale information as the final enhanced feature. Among them, the multi-path parallel dilated convolution branches use the same size convolution kernel and set different dilation rates to form multiple receptive fields from local fineness to global context.
[0019] A further improvement of the technical solution of this invention lies in that the input selected layer features are subjected to global max pooling and global average pooling to extract channel-level global context descriptions, and the results of the two pooling operations, global max pooling and global average pooling, are concatenated to obtain global features. The expression for the global features is as follows:
[0020]
[0021] In the formula, Represents the global characteristics of the corresponding channel; GAP(·) represents global average pooling operation; GMP(·) represents global max pooling operation; This represents element-wise addition; A layer of features representing the input;
[0022] The process involves inputting global features into a two-layer fully connected network. The first layer uses ReLU activation to capture non-linear channel dependencies, and the second layer uses Sigmoid activation to generate normalized channel attention weights. These channel attention weights are then multiplied channel-by-channel with the selected layer features to obtain weighted features. The expression for these weighted features is as follows:
[0023]
[0024] In the formula, Represents a weighted feature; w c Represents channel attention weight; ε(·) represents element-wise multiplication; ε(·) represents the sigmoid activation function; fc1(·) represents the first fully connected layer, fc2(·) represents the second fully connected layer; δ(·) represents the ReLU activation function.
[0025] A further improvement of the technical solution of the present invention is that, in the process of the feature extraction module inputting an image and extracting multi-layer features, the multi-layer features are specifically three-layer features;
[0026] The hierarchical progressive attention fusion module utilizes attention mechanisms and skip connections to fuse and enhance features, reducing information redundancy and offset between features at different levels. The specific steps include:
[0027] The first-layer features and the second-layer features are added element-wise and fused to construct a low-level fused representation. The low-level fused representation is then input into the convolutional attention module to explicitly model the importance of channels and spatial dimensions, generating enhanced attention features. The enhanced attention features are then added element-wise with the first-layer features and the second-layer features to form a transition feature representation.
[0028] The transition feature representation is first fused with the third layer feature by element-wise addition, and then attention enhancement is performed through the convolutional attention module. The attention enhancement result is then added element-wise with the first layer feature, the second layer feature, the third layer feature, and the transition feature, and finally outputs a global enhanced feature containing multi-level collaborative semantics.
[0029] A further improvement of the technical solution of the present invention is that, in the process of the feature extraction module inputting an image and extracting multi-layer features, a ResNet-50 network is used to extract features, and the third-layer feature map, the fourth-layer feature map, and the fifth-layer feature map are retained as outputs; wherein, the output third-layer feature map, the fourth-layer feature map, and the fifth-layer feature map correspond to the first-layer features, the second-layer features, and the third-layer features in the multi-layer features, respectively.
[0030] A further improvement to the technical solution of the present invention lies in the following specific training steps of the cross-modal target tracking model:
[0031] The dataset was constructed as follows: the RGB dataset GOT10K was used as the training dataset for the benchmark algorithm SiamRPN++, the RGBT dataset LasHeR was used as the training set, and GTOT and RGBT234 were used as the test set.
[0032] The model construction includes: modal extension of the baseline network, construction of a multi-scale collaborative module, and construction of a multi-level collaborative module. Specifically, when extending the baseline network, the backbone is first extended to include four branches suitable for RGBT target tracking. These four branches are used for feature extraction and correspond to visible light template branch, visible light search branch, thermal infrared template branch, and thermal infrared search branch, respectively. Furthermore, the original Region Proposal Network (RPN) is extended to an RGBT modal RPN network: depthwise separable correlations are calculated in parallel for the template and search features of visible light and thermal infrared, respectively, outputting two classification score maps and regression offset maps. The results are then fused using soft weights. When constructing the multi-scale collaborative module, multiple multi-scale channel adaptive enhancement modules are designed. When constructing the multi-level collaborative module, a hierarchical progressive attention fusion module is designed.
[0033] The training model includes: First, in the single-modal pre-training stage, the model weights of the baseline algorithm model SiamRPN++ are obtained by using the RGB dataset GOT-10K for visible light single-modal training. Then, the baseline model is extended to RGBT multimodal input, and a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module are added. The LasHeR dataset is used as the RGBT dataset to train the model in stages: first, the backbone parameters are frozen and only other parts are trained; when the training rounds reach the preset number of rounds, the backbone layers are unfrozen and training continues until the model loss value converges. Finally, the complete model is saved to obtain the trained cross-modal target tracking model.
[0034] A second aspect of the present invention provides a cross-modal target tracking system based on a cooperative strategy, comprising:
[0035] The image acquisition module is used to acquire the template image and the search image for the current frame;
[0036] The target tracking module is used to perform cross-modal target tracking based on the acquired template image and the current frame search image, using a trained cross-modal target tracking model to obtain cross-modal target tracking results;
[0037] The cross-modal target tracking model includes four branch networks and a multimodal region proposal network. The four branch networks are respectively used as inputs for a visible light template image, a thermal infrared template image, a visible light search image, and a thermal infrared search image, and respectively output visible light template enhancement features, thermal infrared template enhancement features, visible light search enhancement features, and thermal infrared search enhancement features. The multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates target candidate regions through a fully convolutional network, which serve as the cross-modal target tracking result.
[0038] The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. The extracted features are then processed by the multi-scale channel adaptive enhancement modules through channel attention enhancement and parallel dilated convolution to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are then fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and to obtain the final global enhanced features containing multi-level collaborative semantics.
[0039] In a third aspect, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the cross-modal target tracking method based on a cooperative strategy as described in any one of the first aspects of the present invention.
[0040] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the cross-modal target tracking method based on a cooperative strategy as described in any one of the first aspects of the present invention.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] This invention discloses a cross-modal target tracking method based on a collaborative strategy. This method introduces a multi-scale channel adaptive enhancement module (MCAEM) and a hierarchical progressive attention fusion module (HPAFM) to achieve efficient collaborative fusion of visible light and thermal infrared modal features. This effectively addresses challenges such as scale changes, deformation, and texture blurring, improving tracking stability and robustness under occlusion, modal interference, and low-light conditions. Specifically, addressing the problem of insufficient modal information fusion in existing methods, this invention constructs a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module. By introducing multi-scale structure perception and dynamic attention mechanisms, efficient collaborative fusion between infrared and visible light features is achieved. This design enhances the complementary expressive ability between modalities and strengthens the discriminative power of the fused features, thereby significantly improving tracking accuracy and robustness. To elaborate further, the multi-scale channel adaptive enhancement module can capture the semantic and spatial features of the target at different scales through collaborative modeling with channel attention mechanism and multi-branch dilated convolution, thereby enhancing the perception of key channels and scale regions and effectively addressing challenges such as scale changes, deformation, and texture blurring. In addition, the hierarchical progressive attention fusion module integrates shallow and deep features layer by layer with the help of convolutional attention mechanism, highlights the target region through explicit modeling of channels and space, strengthens multi-layer semantic consistency, and improves the tracking stability and robustness of the network under occlusion, modal interference, and low-light conditions.
[0043] In summary, addressing the lack of scale modeling mechanisms in existing methods, this invention proposes a multi-scale modeling strategy based on dilated convolution. This strategy introduces contextual information from different receptive fields without significantly increasing computational cost, enabling the tracker to adapt to changes in the target such as rapid scaling and local deformation. To address the semantic shift and redundancy issues in existing deep feature fusion methods, this invention introduces a hierarchical progressive fusion strategy. This guides the gradual fusion of deep semantics and shallow details, effectively mitigating inconsistencies in feature representation across different layers and enhancing feature semantic consistency and structural integrity. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0045] Figure 1 This is a flowchart illustrating a cross-modal target tracking method based on a cooperative strategy, as described in an embodiment of the present invention.
[0046] Figure 2 This is a schematic diagram of the entire process of a cross-modal target tracking method based on a cooperative strategy in a specific embodiment of the present invention;
[0047] Figure 3 This is a schematic diagram of the principle architecture of a cross-modal target tracking system in a specific embodiment of the present invention;
[0048] Figure 4 This is a schematic diagram of the structure of the multi-scale channel adaptive enhancement module in an embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the hierarchical attention fusion module in an embodiment of the present invention;
[0050] Figure 6 This is a schematic diagram illustrating the overall performance comparison on the GTOT dataset in an embodiment of the present invention;
[0051] Figure 7 This is a schematic diagram illustrating the overall performance comparison on the RGBT234 dataset in an embodiment of the present invention;
[0052] Figure 8 This is a schematic diagram of a cross-modal target tracking system based on a cooperative strategy, as described in an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention; obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0054] Based on the technical solutions disclosed in the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0055] Please see Figure 1 The present invention provides a cross-modal target tracking method based on a cooperative strategy, comprising the following steps:
[0056] Step 1: Obtain the template image and the current frame search image; further explained, the current frame search image contains the target to be tracked;
[0057] Step 2: Based on the template image obtained in Step 1 and the current frame search image, perform cross-modal target tracking using the trained cross-modal target tracking model to obtain cross-modal target tracking results; in a specific exemplary technical solution, the cross-modal target tracking results include: the coordinates (such as center coordinates or preset vertex coordinates), width, and height of the bounding box of the target to be tracked.
[0058] In this embodiment of the invention, the aforementioned cross-modal target tracking model includes: four branch networks and a multimodal region proposal network; wherein,
[0059] The four branch networks are used as inputs for the visible light template image, the thermal infrared template image, the visible light search image, and the thermal infrared search image, respectively, and output the visible light template enhancement features, the thermal infrared template enhancement features, the visible light search enhancement features, and the thermal infrared search enhancement features, respectively; the multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates high-quality target candidate regions through a fully convolutional network as cross-modal target tracking results;
[0060] The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. The extracted features of each layer are enhanced by channel attention and parallel dilated convolution through the multi-scale channel adaptive enhancement modules to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and obtain the final global enhanced features containing multi-level collaborative semantics.
[0061] In this embodiment of the invention, the multimodal region proposal network includes a visible light branch and a thermal infrared branch. Each branch includes a classification branch and a regression branch. In the classification and regression branches, the input enhancement features are adjusted through convolution operations, and the adjusted template features are used as convolution kernels to perform cross-correlation operations with the adjusted search features to generate corresponding branch response maps (visible light classification branch response map, visible light regression branch response map, thermal infrared classification branch response map, and thermal infrared regression branch response map). Finally, the corresponding branch response maps of visible light and thermal infrared are weighted and fused to obtain the final classification response map and regression response map, and the tracking result is obtained based on the final response map.
[0062] The technical solution disclosed in this invention, while maintaining structural simplicity and operational efficiency, systematically solves key problems such as low fusion quality, scale insensitivity, and semantic offset in existing cross-modal target tracking methods by introducing structural collaborative enhancement and information fusion optimization mechanisms, and has significant performance improvement and wide application value.
[0063] Please see Figure 2 and Figure 3 In a specific embodiment of the present invention, a cross-modal target tracking method based on a cooperative strategy is provided, which mainly includes the following steps:
[0064] (S0) Constructing the dataset, including: using the RGB dataset GOT10K as the training dataset for the benchmark algorithm SiamRPN++, the RGBT dataset LasHeR as the training set, and GTOT and RGBT234 as the test set;
[0065] (S1) Model construction, including: model construction includes modal extension of the baseline network, construction of multi-scale cooperative modules and construction of multi-level cooperative modules;
[0066] (S2) Training the model, including: First, in the single-modal pre-training stage, the model weights of the baseline algorithm model SiamRPN++ are obtained by using the RGB dataset GOT-10K for visible light single-modal training; then, the baseline model is extended to RGBT multimodal input, and a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module are added. The LasHeR dataset is used as the RGBT dataset to train the model in stages: first, the backbone parameters are frozen and only other parts are trained; when the training rounds reach the preset number of rounds, the backbone layers are unfrozen and training continues until the model loss value converges. Finally, the complete model is saved.
[0067] (S3) Obtaining tracking results, including: using the first frame of the visible light and thermal infrared video sequences in the test set as a template image, selecting the target and inputting it into the trained model to generate the final tracking results.
[0068] In this embodiment of the invention, step (S1) is specifically implemented as follows:
[0069] (S1-1) Modal Extension of the Baseline Network: First, the backbone of the baseline network is extended to include four branches suitable for RGBT target tracking, corresponding to the visible light template branch, visible light search branch, thermal infrared template branch, and thermal infrared search branch for feature extraction. Furthermore, to fully utilize the complementarity between deep and shallow features in the target tracking task, this network retains the feature maps from layers three to five during the feature extraction stage. Based on this, the original Region Proposal Network (RPN) is extended to an RGBT modal RPN network: Deeply separable correlations are calculated in parallel for the template and search features of visible light and infrared, respectively, outputting two classification score maps and regression offset maps. The results are then fused using soft weights, thus taking into account the advantages and disadvantages of the two modalities in different scenarios.
[0070] (S1-2) Constructing a multi-scale collaborative module: This invention designs a multi-scale channel adaptive enhancement module, such as... Figure 4 As shown;
[0071] (S1-3) Constructing a multi-level collaborative module: This invention designs a hierarchical progressive attention fusion module, such as... Figure 5 As shown.
[0072] In this embodiment of the invention, step (S3) is specifically implemented as follows:
[0073] (S3-1) Image Acquisition: The template image is centered on the target selected in the first frame of the visible light and thermal infrared image sequence, and cropped to a size of 127×127 pixels. Subsequent frames are search images, with a size of 255×255 pixels;
[0074] (S3-2) Feature Extraction: The two pairs of images obtained from (S3-1) are input into the backbone network of the corresponding branch, and image features at different levels are extracted. Specifically, the third, fourth, and fifth layers of the ResNet-50 network are set as the output layers of the backbone network to obtain features of visible light images and thermal infrared images;
[0075] (S3-3) Feature Enhancement: Multi-scale and multi-level feature enhancement is performed on the features output by (S3-2) using a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module, and the enhanced template features and search features are output.
[0076] (S3-4) Feature Matching: The enhanced features from (S3-3) are adjusted and input into a network containing classification and regression branches for processing. The classification branch calculates the target confidence at each location to determine the target's position; the regression branch performs bounding box regression on the response results to predict the target's precise position and size. In each branch, the algorithm uses the adjusted target features as convolution kernels and performs correlation operations with the feature map of the search image to generate the response map of the corresponding branch. The outputs of the two branches are merged to obtain the position and size of the target in the current frame, completing target tracking.
[0077] Please see Figure 4 In this embodiment of the invention, the multi-scale channel adaptive enhancement module aims to capture and coordinate feature information at different scales using attention mechanisms and dilated convolutions. It mainly comprises two parts: channel attention enhancement and parallel dilated convolutions; wherein,
[0078] First, the multi-layer features extracted from the backbone network are input into a branch. In each branch, the features of that layer first undergo global max pooling and global average pooling to extract channel-level global context descriptions. The specific process is as follows:
[0079]
[0080] in, Representing the global characteristics of the corresponding channel, GAP(·) represents global average pooling operation, and GMP(·) represents global max pooling operation. This represents element-wise addition. This represents a layer of features input to this module.
[0081] Next, the concatenated pooling results are input into a two-layer fully connected network. The first layer uses ReLU activation to capture non-linear channel dependencies, and the second layer uses Sigmoid activation to generate normalized channel attention weights. Finally, the corresponding weights are multiplied channel-by-channel with the input features to obtain the enhanced features. The specific process is as follows:
[0082]
[0083] Among them, w c The values represent the importance weights of the corresponding channels, ε(·) represents the sigmoid activation function, and fc i (·) represents the i-th fully connected layer, and δ(·) represents the ReLU activation function. This represents the global characteristics of the corresponding channel. This represents a layer of features extracted from the backbone network. This represents element-wise multiplication. This represents a feature with weights.
[0084] Interpretationally, channel attention enhancement can automatically identify and amplify the channels that contribute the most to the tracking task, while suppressing irrelevant or noisy channels, thereby enhancing the representational quality of the feature map.
[0085] After channel weighting, a four-way parallel dilated convolution branch is used to extract multi-scale information. All four branches use the same-sized kernel but have different dilation rates, thus forming multiple receptive fields ranging from local fine-grained to global context. Dilated convolution effectively expands the receptive field while maintaining approximately constant computational cost, allowing the model to obtain rich cross-scale features without additional downsampling or pyramid structures. The outputs of each branch are element-wise added along the channel dimension, fusing them into a dynamic perceptual feature map U containing multi-scale information. c This is used for subsequent similarity matching and bounding box regression. The specific process is as follows:
[0086]
[0087] Among them, U c This represents the dynamic sensing feature map output by this module. This represents element-wise addition. Represents a weighted feature, y i (·) represents the i-th holed convolutional layer, i = 1, 2, 3, 4.
[0088] Please see Figure 5 In this embodiment of the invention, the hierarchical progressive attention fusion module aims to use attention mechanisms and skip connections to fuse and enhance feature information at different levels, thereby reducing information redundancy and offset between features at different levels.
[0089] To explain the specific principle, the three-layer feature map enhanced by the preceding modules is processed sequentially as follows: The first and second layer features are fused element-wise to construct a preliminary low-level fused representation. This fusion result is then input into the Convolutional Attention (CBAM) module to explicitly model the importance of channels and spatial dimensions, generating enhanced attention features. Next, this attention feature is again added element-wise to the original features of the first and second layers to form a transition feature representation. Based on this, the module performs the same operation again with the transition feature and the third layer features: first, element-wise fusion is performed, then attention enhancement is performed through the CBAM module, and the enhanced result is fused with the original features of the first, second, and third layers, as well as the transition feature, finally outputting a globally enhanced feature map containing multi-level collaborative semantics. The specific process is as follows:
[0090]
[0091] in, The i-th layer feature represents the input, where i = 1, 2, 3. This represents the fused feature after incorporating the corresponding features. Represents the enhanced attention features, CBAM(·) represents the channel-space joint attention mechanism, and T represents the transition feature. This represents the global enhanced feature map output by this module.
[0092] Explained, the above part uses a layered attention fusion approach to penetrate multiple feature streams, highlighting key areas and suppressing redundant interference.
[0093] In the specific embodiments of the present invention described above, a ResNet-50 backbone network is used to extract multimodal features, parallel dilated convolution is used to enhance multi-scale feature extraction, and a layer-by-layer attention mechanism is used to optimize the fusion of deep and shallow features. This method significantly improves the intermodal complementarity representation ability and target tracking accuracy, enhances robustness to complex scenes such as occlusion and scale changes, and is applicable to fields such as intelligent security and autonomous driving.
[0094] Please see Figure 6 and Figure 7 In specific embodiments of this invention, comparative experiments were conducted with other mainstream RGBT target tracking algorithms on two commonly used datasets in the RGBT target tracking field, GTOT and RGBT234, to demonstrate the superiority of this invention. On the GTOT dataset, the method proposed in this embodiment significantly outperforms existing comparative algorithms in both accuracy and success rate. Its accuracy is 89.4%, 3.3% higher than the second-place algorithm, and its success rate is 72.1%, 3.9% higher than the second-place algorithm. These results demonstrate that the new scheme of this embodiment possesses excellent overall adaptability in tracking tasks with multiple complex factors. On the RGBT234 dataset, despite the higher challenge and more complex scenarios, the method proposed in this embodiment still maintains a leading position in both accuracy and success rate. Its accuracy is 77.3%, 2.3% higher than the second-place algorithm, and its success rate is 53.6%, 0.3% higher than the second-place algorithm. In summary, thanks to the enhanced target representation through multi-scale and multi-level collaborative mechanisms, the model used in this embodiment can maintain continuous perception and discrimination capabilities of the target state under interference conditions.
[0095] The following are embodiments of the apparatus of the present invention, which can be used to execute embodiments of the method of the present invention. For details not disclosed in the apparatus embodiments, please refer to the embodiments of the method of the present invention.
[0096] Please see Figure 8 In this embodiment of the invention, a cross-modal target tracking system based on a cooperative strategy is provided, comprising:
[0097] The image acquisition module is used to acquire the template image and the search image for the current frame;
[0098] The target tracking module is used to perform cross-modal target tracking based on the acquired template image and the current frame search image, using a trained cross-modal target tracking model to obtain cross-modal target tracking results;
[0099] The cross-modal target tracking model includes four branch networks and a multimodal region proposal network. The four branch networks are respectively used as inputs for a visible light template image, a thermal infrared template image, a visible light search image, and a thermal infrared search image, and respectively output visible light template enhancement features, thermal infrared template enhancement features, visible light search enhancement features, and thermal infrared search enhancement features. The multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates target candidate regions through a fully convolutional network, which serve as the cross-modal target tracking result.
[0100] The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. The extracted features are then processed by the multi-scale channel adaptive enhancement modules through channel attention enhancement and parallel dilated convolution to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are then fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and to obtain the final global enhanced features containing multi-level collaborative semantics.
[0101] In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used to execute the operation of a cross-modal target tracking method based on a cooperative strategy.
[0102] In one embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the operating system of the terminal. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM (Random Access Memory) or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the cross-modal target tracking method based on cooperative strategies in the above embodiments.
[0103] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.
[0104] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A cross-modal target tracking method based on a cooperative strategy, characterized in that, Includes the following steps: Get the template image and the search image in the current frame; Based on the acquired template image and the current frame search image, cross-modal target tracking is performed using the trained cross-modal target tracking model to obtain cross-modal target tracking results; The cross-modal target tracking model includes four branch networks and a multimodal region proposal network. The four branch networks are respectively used as inputs for a visible light template image, a thermal infrared template image, a visible light search image, and a thermal infrared search image, and respectively output visible light template enhancement features, thermal infrared template enhancement features, visible light search enhancement features, and thermal infrared search enhancement features. The multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates target candidate regions through a fully convolutional network, which serve as the cross-modal target tracking result. The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. Each extracted layer feature is then processed by the multi-scale channel adaptive enhancement module through channel attention enhancement and parallel dilated convolution to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are then fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and to obtain the final global enhanced features containing multi-level collaborative semantics.
2. The cross-modal target tracking method based on a cooperative strategy according to claim 1, characterized in that, The multimodal region proposal network includes: a visible light branch and a thermal infrared branch; wherein... Both the visible light branch and the thermal infrared branch include classification branches and regression branches; In the classification and regression branches, the input enhancement features are adjusted through convolution operations, and the adjusted template features are used as convolution kernels to perform cross-correlation operations with the adjusted search features to generate visible light classification branch response maps, visible light regression branch response maps, thermal infrared classification branch response maps, and thermal infrared regression branch response maps. The visible light and thermal infrared corresponding branch response maps are weighted and fused to obtain the final classification response map and regression response map, and the cross-modal target tracking results are obtained based on the final response maps.
3. The cross-modal target tracking method based on a cooperative strategy according to claim 1, characterized in that, The multi-scale channel adaptive enhancement module includes the following steps for performing channel attention enhancement and parallel dilated convolution processing to obtain enhanced features that fuse features from different scales: First, the input selected layer features are processed by global max pooling and global average pooling to extract channel-level global context descriptions. The results of the two pooling processes, global max pooling and global average pooling, are then concatenated to obtain the concatenated result. Then, the splicing result is input into a two-layer fully connected network. The first fully connected network uses ReLU activation to capture non-linear channel dependencies, and the second fully connected network uses Sigmoid activation to generate normalized channel attention weights. The channel attention weights are multiplied with the input selected layer features channel by channel to obtain weighted features. Finally, based on the weighted features, multi-path parallel dilated convolution branches are used to extract multi-scale information. The outputs of each branch are added element-wise along the channel dimension and fused into a dynamic perceptual feature map with multi-scale information as the final enhanced feature. Among them, the multi-path parallel dilated convolution branches use the same size convolution kernel and set different dilation rates to form multiple receptive fields from local fineness to global context.
4. The cross-modal target tracking method based on a cooperative strategy according to claim 3, characterized in that, The selected layer features of the input are processed by global max pooling and global average pooling to extract channel-level global context descriptions. The results of the two pooling processes, global max pooling and global average pooling, are then concatenated to obtain the global features. The expression for the global features is as follows: In the formula, Represents the global characteristics of the corresponding channel; GAP(·) represents global average pooling; GMP(·) represents global max pooling; ⊕ represents element-wise addition; A layer of features representing the input; The process involves inputting global features into a two-layer fully connected network. The first layer uses ReLU activation to capture non-linear channel dependencies, and the second layer uses Sigmoid activation to generate normalized channel attention weights. These channel attention weights are then multiplied channel-by-channel with the selected layer features to obtain weighted features. The expression for these weighted features is as follows: In the formula, Represents a feature with weights; w c Represents channel attention weight; This represents element-wise multiplication; ε(·) represents the sigmoid activation function; fc1(·) represents the first fully connected layer, fc2(·) represents the second fully connected layer; δ(·) represents the ReLU activation function.
5. The cross-modal target tracking method based on a cooperative strategy according to claim 1, characterized in that, During the process of the feature extraction module inputting an image and extracting multi-layer features, the multi-layer features specifically consist of three layers of features. The hierarchical progressive attention fusion module utilizes attention mechanisms and skip connections to fuse and enhance features, reducing information redundancy and offset between features at different levels. The specific steps include: The first-layer features and the second-layer features are added element-wise and fused to construct a low-level fused representation. The low-level fused representation is then input into the convolutional attention module to explicitly model the importance of channels and spatial dimensions, generating enhanced attention features. The enhanced attention features are then added element-wise with the first-layer features and the second-layer features to form a transition feature representation. The transition feature representation is first fused with the third layer feature by element-wise addition, and then attention enhancement is performed through the convolutional attention module. The attention enhancement result is then added element-wise with the first layer feature, the second layer feature, the third layer feature, and the transition feature, and finally outputs a global enhanced feature containing multi-level collaborative semantics.
6. The cross-modal target tracking method based on a cooperative strategy according to claim 5, characterized in that, During the process of the feature extraction module inputting an image and extracting multi-layer features, a ResNet-50 network is used to extract features, and the third, fourth, and fifth layer feature maps are retained as outputs; wherein, the output third, fourth, and fifth layer feature maps correspond to the first, second, and third layer features in the multi-layer features, respectively.
7. The cross-modal target tracking method based on a cooperative strategy according to claim 1, characterized in that, The specific training steps for the cross-modal target tracking model are as follows: The dataset was constructed as follows: the RGB dataset GOT10K was used as the training dataset for the benchmark algorithm SiamRPN++, the RGBT dataset LasHeR was used as the training set, and GTOT and RGBT234 were used as the test set. The model construction includes: modal extension of the baseline network, construction of a multi-scale collaborative module, and construction of a multi-level collaborative module. Specifically, when extending the baseline network, the backbone is first extended to include four branches suitable for RGBT target tracking. These four branches are used for feature extraction and correspond to visible light template branch, visible light search branch, thermal infrared template branch, and thermal infrared search branch, respectively. Furthermore, the original Region Proposal Network (RPN) is extended to an RGBT modal RPN network: depthwise separable correlations are calculated in parallel for the template and search features of visible light and thermal infrared, respectively, outputting two classification score maps and regression offset maps. The results are then fused using soft weights. When constructing the multi-scale collaborative module, multiple multi-scale channel adaptive enhancement modules are designed. When constructing the multi-level collaborative module, a hierarchical progressive attention fusion module is designed. The training model includes: First, in the single-modal pre-training stage, the model weights of the baseline algorithm model SiamRPN++ are obtained by using the RGB dataset GOT-10K for visible light single-modal training. Then, the baseline model is extended to RGBT multimodal input, and a multi-scale channel adaptive enhancement module and a hierarchical progressive attention fusion module are added. The LasHeR dataset is used as the RGBT dataset to train the model in stages: first, the backbone parameters are frozen and only other parts are trained; when the training rounds reach the preset number of rounds, the backbone layers are unfrozen and training continues until the model loss value converges. Finally, the complete model is saved to obtain the trained cross-modal target tracking model.
8. A cross-modal target tracking system based on a cooperative strategy, characterized in that, include: The image acquisition module is used to acquire the template image and the search image for the current frame; The target tracking module is used to perform cross-modal target tracking based on the acquired template image and the current frame search image, using a trained cross-modal target tracking model to obtain cross-modal target tracking results; The cross-modal target tracking model includes four branch networks and a multimodal region proposal network. The four branch networks are respectively used as inputs for a visible light template image, a thermal infrared template image, a visible light search image, and a thermal infrared search image, and respectively output visible light template enhancement features, thermal infrared template enhancement features, visible light search enhancement features, and thermal infrared search enhancement features. The multimodal region proposal network is used as input for the enhancement features output by the four branch networks, and generates target candidate regions through a fully convolutional network, which serve as the cross-modal target tracking result. The four branch networks have the same structure, each including: a feature extraction module, a multi-scale channel adaptive enhancement module, and a hierarchical progressive attention fusion module. The feature extraction module takes the image as input and extracts multi-layer features. The extracted features are then processed by the multi-scale channel adaptive enhancement modules through channel attention enhancement and parallel dilated convolution to obtain enhanced features of each layer that fuse features of different scales. The enhanced features of each layer are then fused and enhanced by the hierarchical progressive attention fusion module using attention mechanisms and skip connections to reduce information redundancy and offset between features of different levels, and to obtain the final global enhanced features containing multi-level collaborative semantics.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the cross-modal target tracking method based on a cooperative strategy as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the cross-modal target tracking method based on a cooperative strategy as described in any one of claims 1 to 7.
Citation Information
Patent Citations
RGBT target tracking method based on cross-modal feature self-enhancement and step-by-step fusion
CN118037769A
Cross-modal target tracking method and system based on fully convolutional twin network
CN119540283A