Multimodal object tracking method and system based on temporal modeling and prompt fine-tuning

By constructing a multimodal target tracking network model with a dual-stream architecture and using the spatiotemporal Transformer encoder and temporal cue guide for feature extraction and fusion, the problems of temporal modeling and inter-modal differences in multimodal target tracking are solved, achieving efficient and robust target tracking effects.

CN120510480BActive Publication Date: 2025-10-14XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510996493.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-14
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

Existing multimodal target tracking methods face challenges in temporal modeling and combating inter-modal differences in dynamic scenes. It is difficult to effectively integrate spatiotemporal features, especially when the target is occluded or modally degraded. Existing methods often introduce too many parameters, affecting model efficiency and robustness.

Method used

A multimodal target tracking method based on temporal modeling and cue fine-tuning is adopted to construct a multimodal target tracking network model with a dual-stream architecture, including visible light and infrared light branches, a multimodal temporal cue, a cross-modal intra-frame dual adapter and a bounding box prediction head. Feature extraction and fusion are performed through the spatiotemporal Transformer encoder, temporal cue guide and cross-modal adapter, and only some parameters are updated to adapt to multimodal tasks.

Benefits of technology

Effectively integrate the spatiotemporal features of visible light and infrared images, improve the accuracy and robustness of target tracking, reduce training parameters, enhance the efficiency and effectiveness of multimodal target tracking, and solve the problems of rapid target motion, partial occlusion, and modal degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510480B_ABST
    Figure CN120510480B_ABST
Patent Text Reader

Abstract

The application relates to a kind of multi-modal target tracking method and system based on timing modeling and prompt fine-tuning, the method comprises: constructing multi-modal target tracking network model, including visible light branch, infrared light branch and multi-modal time prompter, cross-modal intra-frame double adapter and boundary box prediction head, each branch includes space-time Transformer encoder and time prompt guide;Space-time Transformer encoder is used for feature extraction, and time prompt guide is used to transfer time information to subsequent frame, and multi-modal time prompter is used to enhance the time prompt of dominant mode, and cross-modal intra-frame double adapter is used to fuse the modal space features of double branch, and boundary box prediction head is used to predict tracking result;The model is trained, and only the parameters of time prompt guide, multi-modal time prompter and cross-modal intra-frame double adapter are updated in the training process;The trained model is applied to multi-modal target tracking.The method and system can improve the accuracy, robustness and efficiency of target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a multimodal target tracking method and system based on temporal modeling and prompt fine-tuning. Background Art

[0002] Multimodal target tracking is an important research direction in the field of computer vision. Its goal is to continuously track targets using information from multiple modalities (such as visible light and infrared light). In real-world scenarios, the effectiveness of single-modal target tracking is easily limited by complex environments such as lighting conditions, target occlusion, and rapid motion. Although existing methods have achieved certain results through multimodal information fusion, they still face many challenges in temporal modeling and combating inter-modal differences in dynamic scenes. Multimodal target tracking not only needs to address the problem of feature differences between modalities, but also needs to address the ability to flexibly adapt to dynamic changes in targets in complex scenes, especially when the target is occluded or modally degraded.

[0003] Traditional multimodal object tracking methods primarily focus on fusing spatial features across modalities. By involving specialized modules and leveraging fully fine-tuned approaches, visible-light-based object tracking models are adapted to multimodal object tracking tasks. Y. Gao et al. designed a deep adaptive fusion network that adaptively integrates modal features in a simple and effective manner through a recursive fusion chain. CL. Li et al. proposed a guidance module that transfers discriminative features from one modality to another to enhance the discriminative power of certain weak modalities.

[0004] These methods introduce many parameters and limit the efficiency and robustness of multimodal tracking models. To address this issue, Yang et al. designed a multimodal visually cued object tracking network. This network transforms multimodal input into a single modality through a cueing paradigm, leveraging the tracking capabilities of a large-scale pre-trained visible light-based model. Building on this, Zhu et al. designed a learnable visually cued multimodal tracking model. By learning modality-dependent cues, the model adapts the frozen pre-trained base model to various downstream multimodal tasks, while introducing only a small number of trainable parameters.

[0005] However, the above methods ignore the temporal contextual relationship in the video and are difficult to cope with changes in realistic dynamic environments. Wang et al. proposed a time-adaptive multimodal tracking framework with a spatiotemporal dual-branch structure. It captures temporal information by updating the template online and comprehensively utilizes spatiotemporal information and multimodal information for target positioning. Sun et al. designed a dual-branch network, introduced independent dynamic template tokens to interact with the search area, embedded temporal information to solve appearance changes, and retained the participation of the initial static template tokens in the joint feature extraction process to ensure that the original and reliable target appearance information is retained to prevent target appearance deviations caused by traditional time updates. However, the existing multimodal target tracking still has limited use of temporal information. In addition, the exploration of the complementarity of spatiotemporal features of different modalities has not been further explored. Summary of the Invention

[0006] The purpose of the present invention is to provide a multimodal target tracking method and system based on temporal modeling and prompt fine-tuning, which can effectively integrate the spatiotemporal features of visible light and infrared images to improve the accuracy, robustness and efficiency of target tracking.

[0007] To achieve the above objectives, the present invention adopts a technical solution: a multimodal target tracking method based on temporal modeling and prompt fine-tuning, comprising the following steps:

[0008] 1) Obtain video sequences in both visible and infrared modalities from a multimodal object tracking dataset. Each modality consists of multiple template image frames and search image frames. These are processed to form a sequence of template image tokens and a sequence of search image tokens, which are then combined with temporal cue tokens to form training data for model training.

[0009] 2) Constructing a multimodal target tracking network model with a dual-stream architecture. The multimodal target tracking network model includes a visible light branch, an infrared light branch, a multimodal temporal cue, a cross-modal intra-frame dual adapter, and a bounding box prediction head. Each branch includes a spatiotemporal Transformer encoder and a temporal cue director. The spatiotemporal Transformer encoder is used for feature extraction. The temporal cue director transfers spatiotemporal features to the bounding box prediction head based on the features extracted by the spatiotemporal Transformer encoder, as well as historical time information to subsequent frames. The multimodal temporal cue enhances the temporal cue token of the dominant modality by fusing the temporal cue tokens of the two modalities. The cross-modal intra-frame dual adapter is used to adaptively fuse modal spatial features between the two branches. The bounding box prediction head predicts tracking results based on the input visual features.

[0010] The multimodal target tracking network model is trained using the training data obtained in step 1) to optimize the model parameters; during the training process, only the parameters of the time cue guide, the multimodal time cue, and the cross-modal intraframe dual adapter are updated;

[0011] 3) Apply the trained multimodal target tracking network model to multimodal target tracking.

[0012] Furthermore, in step 1), L visible light video sequences and corresponding L infrared light video sequences are extracted from the multimodal target tracking dataset using a video-level sampling method. Each visible light video sequence includes N consecutive visible light image frames, and each infrared light video sequence includes N consecutive infrared light image frames. For both visible light and infrared light modalities, the N image frames of each modality are divided into K template image frames and M search image frames.

[0013] Process the K template image frames and M search image frames of each modality to form a template image token sequence and a search image token sequence. Randomly initialize the time cue token of each modality and combine it with the template image token sequence and search image token sequence of the corresponding modality to form the training data of the modality. The training data of the two modalities form a set of training data. For L visible light video sequences and the corresponding L infrared light video sequences, L sets of training data are formed.

[0014] When training a multimodal object tracking network model, for a set of training data input to the model, the randomly initialized temporal cue token, template image token sequence, and the first search image token in the search image token sequence are concatenated as the input sequence for the first iteration. At the tth iteration, the tth search image token is extracted from the search image token sequence and concatenated with the temporal cue token and template image token sequence output at the t-1th iteration as the input sequence for the tth iteration. A complete model training is completed after M iterations, where t = 2, 3, ..., M.

[0015] The multimodal target tracking network model is trained L times using L groups of training data to obtain a trained multimodal target tracking network model.

[0016] Furthermore, in step 2), the multimodal target tracking network model is implemented as follows:

[0017] The multimodal target tracking network model adopts a dual-stream architecture consisting of a visible light branch and an infrared light branch, each of which is composed of a multi-layered spatiotemporal Transformer encoder and a temporal cue guide. The multimodal target tracking network model also includes a multimodal temporal cue, a cross-modal intra-frame dual adapter, and a bounding box prediction head.

[0018] The spatiotemporal Transformer encoder consists of a multi-head attention mechanism and a multi-layer perceptron; the two branches use parameter-sharing spatiotemporal Transformer encoders to extract features from the input sequences of the corresponding modalities;

[0019] The temporal cue guide is based on the features extracted by the spatiotemporal Transformer encoder. It captures the spatiotemporal dependencies between the template image token sequence and the search image token corresponding to the current frame through a spatiotemporal aggregation attention mechanism. It aggregates the obtained new temporal cue token and passes it to the subsequent frame to guide the feature extraction of the subsequent frame. At the same time, the extracted spatiotemporal features are passed to the bounding box prediction head. The temporal cue token is continuously optimized throughout the model training process to adapt to the dynamic scenarios of the target tracking task.

[0020] The multimodal temporal cue generator uses the temporal cues of the auxiliary modality to optimize the temporal cues of the dominant modality, reduces redundant information through convolution operations and spatial saliency operations, and realizes the fusion of spatiotemporal information between frames across modalities, dynamically optimizes the temporal cues of the dominant modality, and generates enhanced temporal cues that are transmitted to subsequent frames of the dominant modality;

[0021] The cross-modal intra-frame dual adapter dynamically adapts to changes in the two modalities through an adapter block fine-tuning mechanism, preventing over-reliance on a single modality and achieving dynamic fusion of intra-frame spatiotemporal information between modalities.

[0022] The bounding box prediction head outputs the target position, offset value and normalized bounding box based on the obtained spatiotemporal features of visible light and infrared light. The spatiotemporal Transformer encoders of the visible light branch and the infrared light branch adopt a pre-trained model and perform parameter fine-tuning during the model training process, that is, all parameters of the pre-trained model are frozen, and only the parameters of the time cue guide, multimodal time cue and cross-modal intra-frame dual adapter are updated to guide the multimodal target tracking network model to perform parameter fine-tuning and adapt to the multimodal target tracking task.

[0023] Furthermore, the temporal cue token is a learnable parameter sequence of length P and dimension d; the temporal cue guide is composed of a spatiotemporal aggregation attention mechanism and a multi-layer perceptron; wherein the spatiotemporal aggregation attention mechanism is used to establish contextual relationships between video sequences and output corresponding feature representations;

[0024] For the visible light branch, it is expressed as follows:

[0025]

[0026] in, represents the attention features aggregated over all spatiotemporal positions in the i-th frame; Attn(·) represents the multi-head attention mechanism for spatiotemporal aggregation, i.e., the spatiotemporal aggregation attention mechanism; represents the i-th frame feature extracted by the spatiotemporal Transformer encoder in the i-th iteration, i = 1, 2, ..., M-1;

[0027] The attention features aggregated from all spatiotemporal positions of the i-th frame are further processed by a multi-head perceptron to output a token sequence, which is divided into a time prompt token, a template image token sequence, and a search image frame token, as shown below:

[0028]

[0029] =split( )

[0030] in, Represents the Token sequence output by the multilayer perceptron, MLP(·) represents the multilayer perceptron operation, split(·) represents the split operation, 、 、 They represent the visible light time cue token, template image token sequence, and search image frame token obtained by the time cue guide of the visible light branch respectively;

[0031] For the infrared light branch, the infrared light time prompt Token is obtained in the same way. , template image Token sequence And search image frame Token ;

[0032] Among them, the search image frame Token of visible light and infrared light search image frame Token The splicing is used as the input feature of the bounding box prediction head, and the tracking result of the i-th frame is predicted by the bounding box prediction head;

[0033] For the Mth iteration, the Mth frame features extracted by the spatiotemporal Transformer encoder are no longer input into the temporal cue guide. Instead, the visible light search image frame token and the infrared light search image frame token are directly extracted from the Mth frame features extracted by the two modalities and spliced ​​as the input features of the bounding box prediction head. The tracking result of the Mth frame is predicted by the bounding box prediction head.

[0034] Furthermore, for the two modes of visible light and infrared light, one is used as the dominant mode and the other as the auxiliary mode;

[0035] For the dominant modality, the search image frame tokens obtained from the two modalities are input into the multimodal time prompter to obtain an enhanced time prompt token. The enhanced time prompt token, the template image token sequence obtained from the modality, and the next search image token of the corresponding modality are then concatenated as the input sequence for the next iteration of the modality.

[0036] For the auxiliary modality, the time prompt token and template image token sequence obtained by the modality are directly concatenated with the next search image token of the corresponding modality as the input sequence for the next iteration of the modality.

[0037] Furthermore, the multimodal time cue tokens are received from different modalities, one being the dominant modality and the other being the auxiliary modality. The dimension of the time cue token is reduced through a 1×1 convolution operation, and the redundant information of the dominant modality is reduced through a spatial saliency operation. The time cue tokens of the dominant and auxiliary modalities are then concatenated and flattened to obtain an enhanced time cue.

[0038] Furthermore, with the visible light time prompt token as the dominant mode, the processing process of the multimodal time prompter is expressed as follows:

[0039]

[0040]

[0041] in, represents the temporal cue token of visible light obtained after the spatial saliency operation, Fovea(·) represents the spatial saliency operation, Conv(·) represents the 1×1 convolution operation, The time hint token of the visible light obtained by the time hint guide of the visible light branch, Indicates the infrared light time prompt token obtained by the infrared light branch time prompt guide, represents the enhanced time hint Token, and Flatten(·) represents the flattening operation.

[0042] Furthermore, the cross-modal intra-frame dual adapter includes two adapter blocks, each consisting of a lower fully connected layer, a GELU, and an upper fully connected layer. The two adapter blocks are inserted between the two branches of the spatiotemporal Transformer encoder, respectively, in parallel with the multi-head attention mechanism and multi-layer perceptron operation in the spatiotemporal Transformer encoder; the parameters of the pre-trained spatiotemporal Transformer encoder are frozen, and only the parameters of the adapter block are updated and optimized;

[0043] In the visible light branch, for the spatiotemporal Transformer encoder of the lth layer, the visible light Token sequence output by the spatiotemporal Transformer encoder of the (l-1)th layer is Perform multi-head attention operation and convert the infrared light Token sequence of the (l-1) layer Visible light token sequence after the adapter block and the (l-1)th layer And the output of the multi-head attention operation is spliced ​​to obtain the token sequence of the intermediate state of the visible light in the lth layer ; Then the Token sequence of the intermediate state of the visible light layer l Perform multi-layer perceptron operation to convert the Token sequence of the intermediate state of the infrared light in the first layer into Token sequence passing through the adapter block and the intermediate state of the lth layer of visible light And the output of the multi-layer perceptron is spliced ​​to obtain the visible light Token sequence of the lth layer ; means as follows:

[0044]

[0045]

[0046] where l = 1, 2, ..., L, L is the total number of layers in the spatiotemporal Transformer encoder, MHA(·) denotes the multi-head attention mechanism, MLP(·) denotes the multi-layer perceptron operation, and BA(·) denotes the adapter block operation.

[0047] Furthermore, during one iterative training process, the search image frame tokens extracted by the two branches are concatenated and input into the bounding box prediction head. The bounding box prediction head outputs the target position, offset value and normalized bounding box through a fully convolutional neural network.

[0048] Then, the parameters of the multimodal target tracking network model are optimized based on the loss function, which includes classification loss, generalized intersection-over-union loss and L1 loss, and is defined as follows:

[0049]

[0050] in, represents the Focal loss function used for classification, represents the IoU loss function for bounding box regression, represents the bounding box regression loss function, and are weighting factors.

[0051] The present invention also provides a multimodal target tracking system based on temporal modeling and prompt fine-tuning, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method can be implemented.

[0052] Compared with the existing technology, the present invention has the following beneficial effects: the present invention provides a multimodal target tracking method and system based on temporal modeling and prompt fine-tuning. The method and system can effectively fuse the spatiotemporal features of visible light and infrared images through temporal modeling and prompt fine-tuning technology to obtain generalized and robust characteristics of the target object. It can not only effectively solve the problems of rapid movement, partial occlusion and modal degradation of the target, but also achieve accurate and robust target tracking in the case of rapid movement, partial occlusion or modal degradation of the target. It can also effectively reduce the modal gap, perform multimodal feature fusion with fewer training parameters, and improve training efficiency and effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a flowchart of an implementation of a multimodal target tracking method based on temporal modeling and prompt fine-tuning provided by an embodiment of the present invention;

[0054] Figure 2 4 is an architecture diagram of a multimodal target tracking network model in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0056] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs.

[0057] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0058] like Figure 1 As shown, this embodiment provides a multimodal target tracking method based on temporal modeling and prompt fine-tuning, including the following steps:

[0059] 1) Obtain video sequences in both visible and infrared modalities from a multimodal object tracking dataset. Each modality consists of multiple template image frames and search image frames. These are processed to form a sequence of template image tokens and a sequence of search image tokens, which are then combined with temporal cue tokens to form training data for model training.

[0060] 2) Constructing a multimodal target tracking network model with a dual-stream architecture. The multimodal target tracking network model includes a visible light branch, an infrared light branch, a multimodal temporal cue, a cross-modal intra-frame dual adapter, and a bounding box prediction head. Each branch includes a spatiotemporal Transformer encoder and a temporal cue director. The spatiotemporal Transformer encoder is used for feature extraction. The temporal cue director transfers spatiotemporal features to the bounding box prediction head based on the features extracted by the spatiotemporal Transformer encoder, as well as historical time information to subsequent frames. The multimodal temporal cue enhances the temporal cue token of the dominant modality by fusing the temporal cue tokens of the two modalities. The cross-modal intra-frame dual adapter is used to adaptively fuse modal spatial features between the two branches with a small number of parameters. The bounding box prediction head predicts tracking results based on the input visual features.

[0061] The multimodal target tracking network model is trained using the training data obtained in step 1) to optimize the model parameters; during the training process, only the parameters of the time cue guide, the multimodal time cue, and the cross-modal intraframe dual adapter are updated;

[0062] 3) Apply the trained multimodal target tracking network model to multimodal target tracking.

[0063] In this embodiment, the multimodal target tracking dataset is the LasHeR dataset.

[0064] In step 1), video-level sampling (data extraction in units of video sequences) is used to extract L visible light video sequences and corresponding L infrared light video sequences from the multimodal target tracking dataset. Each visible light video sequence includes N consecutive visible light image frames, and each infrared light video sequence includes N consecutive infrared light image frames. For both visible light and infrared light modalities, the N image frames of each modality are divided into K template image frames and M search image frames.

[0065] The K template image frames and M search image frames of each modality are processed to form a template image Token sequence and a search image Token sequence. The time cue Token of each modality is randomly initialized and combined with the template image Token sequence and the search image Token sequence of the corresponding modality to form the training data of the modality. The training data of the two modalities form a set of training data. For L visible light video sequences and the corresponding L infrared light video sequences, L sets of training data are formed.

[0066] When training a multimodal object tracking network model, for a set of training data input to the model, the randomly initialized temporal cue token, template image token sequence, and the first search image token in the search image token sequence are concatenated as the input sequence for the first iteration. During the t-th iteration, the t-th search image token is extracted from the search image token sequence and concatenated with the temporal cue token and template image token sequence output at the t-1-th iteration as the input sequence for the t-th iteration. A complete model training is completed after M iterations, where t = 2, 3, ..., M.

[0067] The multimodal target tracking network model is trained L times using L groups of training data to obtain a trained multimodal target tracking network model.

[0068] Figure 2 This is the architecture diagram of the multimodal target tracking network model in this embodiment. Figure 2 As shown, in step 2), the implementation method of the multimodal target tracking network model is specifically as follows.

[0069] The multimodal target tracking network model is a dual-stream architecture, which consists of a visible light branch and an infrared light branch. Each branch is composed of a multi-layer stacked spatiotemporal Transformer encoder and a time cue guide; the multimodal target tracking network model also includes a multimodal time cue, a cross-modal intra-frame dual adapter and a bounding box prediction head.

[0070] The spatiotemporal Transformer encoder consists of a multi-head attention mechanism and a multi-layer perceptron; the two branches use a parameter-sharing spatiotemporal Transformer encoder to extract features from the input sequence of the corresponding modality.

[0071] The temporal cue guide is based on the features extracted by the spatiotemporal Transformer encoder. It captures the spatiotemporal dependency between the template image Token sequence and the search image Token corresponding to the current frame through a spatiotemporal aggregation attention mechanism, aggregates the obtained new temporal cue Token and passes it to the subsequent frames to guide the feature extraction of the subsequent frames. At the same time, the extracted spatiotemporal features are passed to the bounding box prediction head. The temporal cue Token is continuously optimized throughout the model training process to adapt to the dynamic scenarios of the target tracking task.

[0072] The multimodal time prompter utilizes the time prompts of the auxiliary modality to optimize the time prompts of the dominant modality, reduces redundant information through convolution operations and spatial saliency operations, and realizes the fusion of spatiotemporal information between frames across modalities, dynamically optimizes the time prompts of the dominant modality, and generates enhanced time prompts to be transmitted to subsequent frames of the dominant modality.

[0073] The cross-modal intra-frame dual adapter dynamically adapts to changes in the two modalities through an adapter block fine-tuning mechanism, prevents over-reliance on a single modality, and achieves dynamic fusion of intra-frame spatiotemporal information between modalities.

[0074] Based on the obtained spatiotemporal features of visible light and infrared light, the bounding box prediction head outputs a target classification score map (i.e., target position), offset value, and normalized bounding box. The spatiotemporal Transformer encoders of the visible light branch and the infrared light branch use the single-modal target tracking model ODTrack as a pre-training model, expanding the single-stream architecture of the original pre-trained model to a dual-stream architecture. Parameters are fine-tuned during model training, that is, all parameters of the pre-trained model are frozen, and only the parameters of the time cue guide, multimodal time cue, and cross-modal intra-frame dual adapter are updated to guide the multimodal target tracking network model to perform parameter fine-tuning and adapt to the multimodal target tracking task.

[0075] In this embodiment, the specific implementation method of the time prompt guide is as follows.

[0076] The time cue token is a learnable parameter sequence with a length of P and a dimension of d. The time cue guide consists of a spatiotemporal aggregation attention mechanism and a multi-layer perceptron. The spatiotemporal aggregation attention mechanism is used to establish contextual relationships between video sequences and output corresponding feature representations.

[0077] Taking visible light as an example, it is expressed as follows:

[0078]

[0079] in, represents the attention features aggregated over all spatiotemporal positions in the i-th frame; Attn(·) represents the multi-head attention mechanism for spatiotemporal aggregation, i.e., the spatiotemporal aggregation attention mechanism; represents the i-th frame feature extracted by the spatiotemporal Transformer encoder in the i-th iteration, i = 1, 2, ..., M-1.

[0080] The attention features aggregated from all spatiotemporal positions of the i-th frame are further processed by a multi-head perceptron to output a token sequence, which is divided into a time prompt token, a template image token sequence, and a search image frame token, as shown below:

[0081]

[0082] =split( )

[0083] in, Represents the Token sequence output by the multilayer perceptron, MLP(·) represents the multilayer perceptron operation, split(·) represents the split operation, 、 、 They respectively represent the visible light time cue Token, template image Token sequence, and search image frame Token obtained by the time cue guide of the visible light branch.

[0084] For the infrared light branch, the infrared light time prompt Token is obtained in the same way. , template image Token sequence And search image frame Token .

[0085] Among them, the search image frame Token of visible light and infrared light search image frame Token The splicing is used as the input feature of the bounding box prediction head, and the tracking result of the i-th frame is predicted by the bounding box prediction head.

[0086] For the Mth iteration, the Mth frame features extracted by the spatiotemporal Transformer encoder are no longer input into the temporal cue guide. Instead, the visible light search image frame token and the infrared light search image frame token are directly extracted from the Mth frame features extracted by the two modalities and spliced ​​as the input features of the bounding box prediction head. The tracking result of the Mth frame is predicted by the bounding box prediction head.

[0087] For the visible light and infrared light modalities, one serves as the dominant modality and the other as the auxiliary modality. For the dominant modality, the search image frame tokens obtained from both modalities are input into the multimodal time cue generator to obtain an enhanced time cue token. The enhanced time cue token, the template image token sequence obtained from the modality, and the next search image token of the corresponding modality are then concatenated as the input sequence for the next iteration of the modality. For the auxiliary modality, the time cue token and template image token sequence obtained from the modality are directly concatenated with the next search image token of the corresponding modality as the input sequence for the next iteration of the modality.

[0088] In this embodiment, the specific implementation method of the multimodal time prompter is as follows.

[0089] The multimodal time cueing device accepts time cue tokens from different modalities, one of which is the dominant modality and the other is the auxiliary modality. The dimension of the time cue token is reduced through a 1×1 convolution operation, and the redundant information of the dominant modality is reduced through a spatial saliency operation. The time cue tokens of the dominant modality and the auxiliary modality are then concatenated and flattened to obtain an enhanced time cue.

[0090] In this embodiment, the visible light time reminder token is used as the dominant mode, and the processing process of the multimodal time reminder is as follows:

[0091]

[0092]

[0093] in, represents the temporal cue token of visible light obtained after the spatial saliency operation, Fovea(·) represents the spatial saliency operation, Conv(·) represents the 1×1 convolution operation, The time hint token of the visible light obtained by the time hint guide of the visible light branch, Indicates the infrared light time prompt token obtained by the infrared light branch time prompt guide, represents the enhanced time hint Token, and Flatten(·) represents the flattening operation.

[0094] In this embodiment, a specific implementation method of the cross-modal intra-frame dual adapter is as follows.

[0095] The cross-modal intra-frame dual adapter includes two adapter blocks, each of which is composed of a lower fully connected layer, a GELU, and an upper fully connected layer. The two adapter blocks are respectively inserted between the spatiotemporal Transformer encoders of the two branches, and operate in parallel with the multi-head attention mechanism and multi-layer perceptron in the spatiotemporal Transformer encoder; the parameters of the pre-trained spatiotemporal Transformer encoder are frozen, and only the parameters of the adapter blocks are updated and optimized.

[0096] In the visible light branch, for the spatiotemporal Transformer encoder of the lth layer, the visible light Token sequence output by the spatiotemporal Transformer encoder of the (l-1)th layer is Perform multi-head attention operation and convert the infrared light Token sequence of the (l-1) layer Visible light token sequence after the adapter block and the (l-1)th layer And the output of the multi-head attention operation is spliced ​​to obtain the token sequence of the intermediate state of the visible light in the lth layer ; Then the Token sequence of the intermediate state of the visible light layer l Perform multi-layer perceptron operation to convert the Token sequence of the intermediate state of the infrared light in the first layer into Token sequence passing through the adapter block and the intermediate state of the lth layer of visible light And the output of the multi-layer perceptron is spliced ​​to obtain the visible light Token sequence of the lth layer ; means as follows:

[0097]

[0098]

[0099] where l = 1, 2, ..., L, L is the total number of layers in the spatiotemporal Transformer encoder, MHA(·) denotes the multi-head attention mechanism, MLP(·) denotes the multi-layer perceptron operation, and BA(·) denotes the adapter block operation.

[0100] During an iterative training process, the search image frame tokens extracted by the two branches are spliced ​​and input into the bounding box prediction head. The bounding box prediction head outputs the target classification score map (i.e., target position) offset value and normalized bounding box through a fully convolutional neural network.

[0101] Then, the parameters of the multimodal target tracking network model are optimized based on the loss function, which includes classification loss, generalized intersection-over-union loss and L1 loss, and is defined as follows:

[0102]

[0103] in, represents the Focal loss function used for classification, represents the general IoU loss function for bounding box regression, represents the bounding box regression loss function, and are weighting factors.

[0104] In this example, the LasHeR dataset was used for comparative verification using visible light and infrared images of the target object. Table 1 shows the comparison results of the proposed method with other multimodal target tracking methods on the LasHeR dataset. As can be seen from Table 1, the proposed method has higher precision and robustness than other multimodal target tracking methods, specifically, achieving the best precision and accuracy.

[0105] Table 1

[0106]

[0107] This embodiment also provides a multimodal target tracking system based on timing modeling and prompt fine-tuning, including a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the above-mentioned method can be implemented.

[0108] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0110] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other manner. Any person skilled in the art may utilize the above-disclosed technical content to modify or modify the present invention into equivalent embodiments. However, any simple modifications, equivalent variations, and modifications to the above embodiments that do not depart from the technical content of the present invention and are based on the technical essence of the present invention remain within the scope of protection of the present invention.

Claims

1. A multimodal target tracking method based on temporal modeling and cue fine-tuning, characterized in that: The following steps are involved: 1) Obtain video sequences in both visible and infrared modalities from a multimodal target tracking dataset. Each modality consists of multiple template image frames and search image frames. These are processed to form a sequence of template image tokens and a sequence of search image tokens, which are then combined with temporal cue tokens to form training data for model training. 2) Constructing a multimodal target tracking network model with a dual-stream architecture, the multimodal target tracking network model includes a visible light branch, an infrared light branch, a multimodal time cue, a cross-modal intra-frame dual adapter, and a bounding box prediction head. Each branch includes a spatiotemporal Transformer encoder and a time cue director. The spatiotemporal Transformer encoder is used for feature extraction. The time cue director transfers spatiotemporal features to the bounding box prediction head based on the features extracted by the spatiotemporal Transformer encoder, and transfers historical time information to subsequent frames. The multimodal time cue enhances the time cue token of the dominant modality by fusing the time cue tokens of the two modalities. The cross-modal intra-frame dual adapter is used to adaptively fuse modal spatial features between the two branches. The bounding box prediction head predicts tracking results based on the input visual features. The multimodal target tracking network model is trained using the training data obtained in step 1) to optimize the model parameters; during the training process, only the parameters of the time prompt director, the multimodal time prompter, and the cross-modal intraframe dual adapter are updated; 3) Apply the trained multimodal target tracking network model to multimodal target tracking; Taking the visible light time prompt token as the dominant mode, the processing process of the multimodal time prompter is as follows: in, represents the temporal cue token of visible light obtained after the spatial saliency operation, Fovea(·) represents the spatial saliency operation, Conv(·) represents the 1×1 convolution operation, The time hint token of the visible light obtained by the time hint guide of the visible light branch, Indicates the infrared light time prompt token obtained by the infrared light branch time prompt guide, represents the enhanced time hint token, and Flatten(·) represents the flattening operation; The cross-modal intra-frame dual adapter includes two adapter blocks, each of which is composed of a lower fully connected layer, a GELU, and an upper fully connected layer. The two adapter blocks are respectively inserted between the spatiotemporal Transformer encoders of the two branches, and operate in parallel with the multi-head attention mechanism and multi-layer perceptron in the spatiotemporal Transformer encoder; the parameters of the pre-trained spatiotemporal Transformer encoder are frozen, and only the parameters of the adapter blocks are updated and optimized.

2. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 1 is characterized in that: In step 1), L visible light video sequences and corresponding L infrared light video sequences are extracted from the multimodal target tracking dataset using video-level sampling. Each visible light video sequence includes N consecutive visible light image frames, and each infrared light video sequence includes N consecutive infrared light image frames. For both visible light and infrared light modalities, the N image frames of each modality are divided into K template image frames and M search image frames. Process the K template image frames and M search image frames of each modality to form a template image token sequence and a search image token sequence. Randomly initialize the time cue token of each modality and combine it with the template image token sequence and search image token sequence of the corresponding modality to form the training data of the modality. The training data of the two modalities form a set of training data. For L visible light video sequences and the corresponding L infrared light video sequences, L sets of training data are formed. When training a multimodal target tracking network model, for a set of training data input to the model, the randomly initialized temporal cue token, template image token sequence, and the first search image token in the search image token sequence are concatenated as the input sequence for the first iteration; when performing the tth iteration, the tth search image token is extracted from the search image token sequence and concatenated with the temporal cue token and template image token sequence output at the t-1th iteration as the input sequence for the tth iteration; M iterations are performed to complete a complete model training; where t = 2, 3, ..., M; The multimodal target tracking network model is trained L times using L groups of training data to obtain a trained multimodal target tracking network model.

3. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 1 is characterized in that: In step 2), the implementation method of the multimodal target tracking network model is: The multimodal target tracking network model adopts a dual-stream architecture consisting of a visible light branch and an infrared light branch, each of which is composed of a multi-layered spatiotemporal Transformer encoder and a temporal cue guide. The multimodal target tracking network model also includes a multimodal temporal cue, a cross-modal intra-frame dual adapter, and a bounding box prediction head. The spatiotemporal Transformer encoder consists of a multi-head attention mechanism and a multi-layer perceptron; the two branches use parameter-sharing spatiotemporal Transformer encoders to extract features from the input sequences of the corresponding modalities; The temporal cue guide is based on the features extracted by the spatiotemporal Transformer encoder. It captures the spatiotemporal dependencies between the template image token sequence and the search image token corresponding to the current frame through a spatiotemporal aggregation attention mechanism. It aggregates the obtained new temporal cue token and passes it to the subsequent frame to guide the feature extraction of the subsequent frame. At the same time, the extracted spatiotemporal features are passed to the bounding box prediction head. The temporal cue token is continuously optimized throughout the model training process to adapt to the dynamic scenarios of the target tracking task. The multimodal temporal cue generator uses the temporal cues of the auxiliary modality to optimize the temporal cues of the dominant modality, reduces redundant information through convolution operations and spatial saliency operations, and realizes the fusion of spatiotemporal information between frames across modalities, dynamically optimizes the temporal cues of the dominant modality, and generates enhanced temporal cues that are transmitted to subsequent frames of the dominant modality; The cross-modal intra-frame dual adapter dynamically adapts to changes in the two modalities through an adapter block fine-tuning mechanism, preventing over-reliance on a single modality and achieving dynamic fusion of intra-frame spatiotemporal information between modalities. The bounding box prediction head outputs the target position, offset value and normalized bounding box based on the obtained spatiotemporal features of visible light and infrared light. The spatiotemporal Transformer encoders of the visible light branch and the infrared light branch adopt a pre-trained model and perform parameter fine-tuning during the model training process, that is, all parameters of the pre-trained model are frozen, and only the parameters of the time cue guide, multimodal time cue and cross-modal intra-frame dual adapter are updated to guide the multimodal target tracking network model to perform parameter fine-tuning and adapt to the multimodal target tracking task.

4. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 3 is characterized in that: The temporal cue token is a learnable parameter sequence of length P and dimension d. The temporal cue guide consists of a spatiotemporal attention mechanism and a multi-layer perceptron. The spatiotemporal attention mechanism is used to establish contextual relationships between video sequences and output corresponding feature representations. For the visible light branch, it is expressed as follows: in, represents the attention features aggregated over all spatiotemporal positions in the i-th frame; Attn(·) represents the multi-head attention mechanism for spatiotemporal aggregation, i.e., the spatiotemporal aggregation attention mechanism; represents the i-th frame feature extracted by the spatiotemporal Transformer encoder in the i-th iteration, i = 1, 2, ..., M-1; The attention features aggregated from all spatiotemporal positions of the i-th frame are subjected to a multi-head perceptron operation to output a token sequence, which is divided into a time prompt token, a template image token sequence, and a search image frame token, as shown below: in, Represents the Token sequence output by the multilayer perceptron, MLP(·) represents the multilayer perceptron operation, split(·) represents the split operation, They represent the visible light time cue token, template image token sequence, and search image frame token obtained by the time cue guide of the visible light branch respectively; For the infrared light branch, the infrared light time prompt Token is obtained in the same way. Template image token sequence And search image frame Token Among them, the search image frame Token of visible light and infrared light search image frame Token The splicing is used as the input feature of the bounding box prediction head, and the tracking result of the i-th frame is predicted by the bounding box prediction head; For the Mth iteration, the Mth frame features extracted by the spatiotemporal Transformer encoder are no longer input into the temporal cue guide. Instead, the visible light search image frame token and the infrared light search image frame token are directly extracted from the Mth frame features extracted by the two modalities and spliced ​​as the input features of the bounding box prediction head. The tracking result of the Mth frame is predicted by the bounding box prediction head.

5. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 4 is characterized in that: For the two modes of visible light and infrared light, one is used as the dominant mode and the other as the auxiliary mode; For the dominant modality, the search image frame tokens obtained from the two modalities are input into the multimodal time prompter to obtain an enhanced time prompt token. The enhanced time prompt token, the template image token sequence obtained from the modality, and the next search image token of the corresponding modality are then concatenated as the input sequence for the next iteration of the modality. For the auxiliary modality, the time prompt token and template image token sequence obtained by the modality are directly concatenated with the next search image token of the corresponding modality as the input sequence for the next iteration of the modality.

6. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 3, characterized in that: The multimodal time cueing device accepts time cue tokens from different modalities, one of which is the dominant modality and the other is the auxiliary modality. The dimension of the time cue token is reduced through a 1×1 convolution operation, and the redundant information of the dominant modality is reduced through a spatial saliency operation. The time cue tokens of the dominant modality and the auxiliary modality are then concatenated and flattened to obtain an enhanced time cue.

7. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 3 is characterized in that: In the visible light branch, for the spatiotemporal Transformer encoder of the lth layer, the visible light Token sequence output by the spatiotemporal Transformer encoder of the (l-1)th layer is Perform multi-head attention operation and convert the infrared light Token sequence of the (l-1) layer Visible light token sequence after the adapter block and the (l-1)th layer And the output of the multi-head attention operation is spliced ​​to obtain the Token sequence of the intermediate state of the visible light in the lth layer Then the Token sequence of the intermediate state of the visible light layer l Perform multi-layer perceptron operation to convert the Token sequence of the intermediate state of the infrared light in the first layer into Token sequence passing through the adapter block and the intermediate state of the lth layer of visible light And the output of the multi-layer perceptron is spliced ​​to obtain the visible light Token sequence of the lth layer It is expressed as follows: where l = 1, 2, ..., L, L is the total number of layers in the spatiotemporal Transformer encoder, MHA(·) denotes the multi-head attention mechanism, MLP(·) denotes the multi-layer perceptron operation, and BA(·) denotes the adapter block operation.

8. The multimodal target tracking method based on temporal modeling and prompt fine-tuning according to claim 3, characterized in that: During one iterative training process, the search image frame tokens extracted by the two branches are spliced ​​and input into the bounding box prediction head. The bounding box prediction head outputs the target position, offset value and normalized bounding box through a fully convolutional neural network. Then, the parameters of the multimodal target tracking network model are optimized based on the loss function, which includes classification loss, generalized intersection-over-union loss and L1 loss, and is defined as follows: L total =L cls +a IoU L IoU +α1L1 Among them, L cls Represents the Focal loss function for classification, L Iou represents the IoU loss function for bounding box regression, L1 represents the L1 loss function for bounding box regression, α IoU and a1 are weighting factors.

9. A multimodal target tracking system based on temporal modeling and cue fine-tuning, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 8 can be implemented.

Citation Information

Patent Citations

  • RGBT target tracking method based on medium-term fusion element framework and composite visual prompt

    CN118967747A

  • Visible light and infrared image fusion target tracking method based on visual prompt learning

    CN119888569A