A multi-modal unified object tracking method based on modal unified representation
Through the combined use of multimodal embedding layer and Transformer model, the redundancy and complexity problems caused by modal independence in the multimodal target tracking method are solved, and the modal unified target tracking model is realized, which improves the flexibility and accuracy of the model.
Patent Information
- Application Number
- CN202510193101.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Due to the heterogeneity of the data characteristics of each modal, the existing multimodal target tracking methods have limited redundant training processes, complex model structures and cross-modal knowledge sharing, making it difficult to process multiple modal inputs at the same time, limiting the algorithm performance and application scope.
By introducing a multimodal embedding layer, multiple modal data such as visible light, depth, infrared, events, and natural language are converted into a unified marking form, and a Transformer model is trained for joint feature extraction and fusion, and task recognition training strategies and soft mark type embedding are used to enhance model performance.
It realizes architecture, model and knowledge sharing of different modal tracking tasks, improves the flexibility and accuracy of the target tracking model, can effectively process five multimodal input signals, and expands the application scope of the algorithm.
Smart Images

Figure CN119672071B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of machine learning, computer vision, and object tracking, and relates to multimodal signal representation algorithms, multi-task joint training algorithms, and multimodal object tracking algorithms; specifically, it is a multimodal unified object tracking method based on unified modal representation. Background Art
[0002] Visual object tracking aims to continuously track an object of interest specified by the user in a continuous video using visual features and temporal information. As one of the basic technologies of computer vision systems, object tracking is widely used in fields such as national defense and military, intelligent transportation, autonomous driving, security monitoring, human-computer interaction, and intelligent commerce.
[0003] Currently, single-modal visual object tracking algorithms based on visible light images have achieved good performance in various scenarios. However, in more complex or specific scenarios, visible light images may not provide sufficient information to support accurate tracking. To address this challenge, many tracking methods have started to use multimodal signals to enhance tracking performance.
[0004] Specifically, depth modal data provides three-dimensional information about the shape and position of the scene and the object, thereby enhancing the robustness of the tracking method to handle occlusion and complex backgrounds. Representative methods include SPT, DeT, DAL, and CA3DMS, etc. Infrared modal data can provide object and background features that are not affected by environmental lighting conditions, and thus is crucial for accurate tracking under low illumination or extreme weather conditions. Representative methods include APFNet, CMPP, JMMAC, and DAFNet, etc. Event modal data is captured at a high frame rate, providing high dynamic range and low-latency perception capabilities, and is suitable for fast-moving or high-speed changing scenarios. Representative methods include OSTrack_E and SiamRCNN_E, etc. Natural language modal data can describe the attributes and behaviors of the object through natural language, further enhancing the understanding and reasoning capabilities of the tracker. Representative methods include JointNLT, DecoupleTNL, VLT TT , and CapsuleTN, etc.
[0005] The above multimodal tracking methods for single tasks have achieved good performance. However, due to the different characteristics of various modal data, these methods design targeted model architectures for each modality and train independent models. This independence between multimodal tracking methods leads to redundant training processes, complex model structures, and limits the sharing of cross-modal knowledge.
[0006] Therefore, how to design and develop an object tracking method that can simultaneously process multiple modal inputs to improve the performance and application scope of the algorithm is one of the urgent challenges in the current tracking field. Summary of the Invention
[0007] The present invention aims to provide a multi-modal unified object tracking method based on modal unified representation. Through a multi-modal embedding layer, visible light, depth, infrared, event, and natural language modalities are represented in a unified token form, making it possible to train a Transformer model for joint feature extraction and fusion of multiple modalities, thereby developing an object tracking model that can handle different multi-modal input signals. In addition, a task recognition training strategy is introduced in this method to enhance the model's ability to distinguish different modal tracking tasks, and soft token type embedding is proposed to provide the model with accurate token type information, further improving the performance of the multi-modal unified model. Finally, this method solves different multi-modal tracking tasks through a unified scheme, achieving architecture unity, model unity, and knowledge sharing among different tasks, and obtaining good tracking performance on five multi-modal tracking tasks (visible light, visible light-depth, visible light-infrared, visible light-event, visible light-natural language tracking).
[0008] Technical solution of the present invention:
[0009] A multi-modal unified object tracking method based on modal unified representation, the steps are as follows:
[0010] Step 1: Obtain videos of multi-modal object tracking tasks from a public dataset for training a multi-modal object tracking model. The videos of multi-modal object tracking tasks include five modalities: visible light, visible light-depth tracking, visible light-infrared tracking, visible light-event tracking, and visible light-natural language. Among them, the depth, infrared, event, and natural language modalities are paired with the visible light modality, and the depth, infrared, and event modalities are stored in the form of three-channel images. Randomly collect image frames in the videos of multi-modal object tracking tasks, and expand the bounding boxes of the interested objects annotated on the image frames by 2 times and 4 times respectively to generate sample pairs of template images and search region images, and perform augmentation using brightness transformation and inversion;
[0011] Step 2: Convert the input signals of different modalities of each multi-modal object tracking task into a unified token embedding form, that is, the unified representation of the modality;
[0012] (1) For visible light-depth object tracking tasks, visible light-infrared object tracking tasks, and visible light-event object tracking tasks, the depth, infrared, and event modalities are collectively referred to as auxiliary modal data and marked as DTE; the visible light modal data is marked as RGB; the visible light modal data is paired with the auxiliary modal data. After concatenating the visible light modal data and the auxiliary modal data in the channel direction, a multi-modal embedding layer is used to jointly perform the conversion of token embedding, so that the visible light modal data and the auxiliary modal data are converted into a unified token embedding representation;
[0013] Concatenate the image and the image along the channel direction to obtain a concatenated image , as shown in the following formula:
[0014]
[0015] Next, the concatenated image is divided into image patches of a fixed size, and the size of each image patch is ; then, each image patch is flattened into a one-dimensional vector with a length of ; finally, a linear transformation is applied to map the flattened image patch vector to the embedding space, as shown in the following formula:
[0016]
[0017] where represents the embedding vector of the th image patch, and its dimension is ; represents the flattened vector of the th image patch, is a weight matrix with a dimension of , is a bias term with a dimension of ;
[0018] (2) For the object tracking task that does not contain DTE data, create a six-channel input by copying the three channels of the RGB data, and then process it using a multi-modal embedding layer to obtain a multi-modal embedding;
[0019] In the visible light-natural language object tracking task, for the natural language modality, use a language model as a text encoder to extract a language feature embedding; among them, the language model is the CLIP-L model, and a linear layer is added to it to adjust the dimension; this language feature embedding is then concatenated with the multi-modal embedding and input into the Transformer encoder;
[0020] (3) For the object tracking task that does not contain the language modality, fill it with a fixed meaningless sentence;
[0021] In the above way, obtain the unified token embeddings of the multi-modal search region, template, and text description; these unified token embeddings are directly fed into a Transformer encoder after being concatenated in the spatial dimension; the internal attention mechanism of the Transformer encoder completes the joint feature extraction and fusion of these unified token embeddings;
[0022] Step 3: Add soft label type encodings to the obtained unified label embeddings to enhance the precise discrimination of the foreground labels of the template image, the background labels of the template image, and the labels of the search region image by the multi-modal tracking model; the foreground labels and background labels of the template image are distinguished by a given bounding box as follows:
[0023] Given a template image containing the target and its bounding box , the template image containing the target is the result of the channel-wise concatenation of the template image obtained in Step 1 after Step 2; first, create a mask with the same size as the template image , in this mask, the pixels inside the bounding box are assigned a value of 1, and the pixels outside the bounding box are assigned a value of 0:
[0024]
[0025] Next, divide the mask into non-overlapping image patches of size ; the th image patch is denoted as ; then, for the values in each image patch, calculate the average value:
[0026]
[0027] where is the average value of the th image patch, representing the degree to which the label embedding corresponding to this image patch is regarded as the foreground;
[0028] Each label type corresponds to a learnable label type embedding, including the foreground label type embedding of the template image, the background label type embedding of the template, and the label type embedding of the search region image. These label type embeddings are learned during the training phase and fixed during the inference phase; weight the label type embeddings corresponding to different types of labels based on the average value of the image patches to enhance the multi-modal image patch embeddings of the template image and the search region image; for the k-th image patch embedding of the template image, make the following adjustment:
[0029]
[0030] where represents the adjusted embedding of the th image patch, represents the original multi-modal image patch embedding, is the foreground label type embedding, It is the background marker type embedding; for the search area image, only the search area marker type embedding is added to each image patch embedding, without distinguishing foreground and background anymore:
[0031]
[0032] Step 4: Utilize the above-mentioned multi-modal marker embedding to train a multi-modal object tracking model; the multi-modal object tracking model consists of a Transformer encoder and a tracking head prediction network; the Transformer encoder is used for feature extraction and fusion of the input multi-modal search area image and template image, adopting the HiViT structure; the tracking head prediction network is used to predict the tracking result on the features output by the Transformer encoder, adopting the head network structure of OSTrack; in order to train the multi-modal object tracking model, the following training and optimization strategies are adopted:
[0033] Adopt a multi-task data mixing training method, that is, mix data from five multi-modal object tracking tasks in each training batch; for the multi-modal object tracking task, use weighted focal loss for foreground and background classification supervision; for the supervision of bounding box regression, adopt a combination of loss and generalized intersection over union loss for supervision; in addition, in order to enhance the ability of the multi-modal object tracking model to distinguish different modal object tracking tasks and corresponding modalities, a task recognition training strategy is also designed; this task recognition training strategy is to make the multi-modal object tracking model explicitly identify the current object tracking task being executed during the training process through an explicit task recognition mechanism, so as to optimize its performance under different object tracking tasks;
[0034] First, take the average calculation of all the feature embeddings output by the Transformer encoder to generate a single feature vector ; the calculation formula of this feature vector is:
[0035]
[0036] where represents the number of output feature embeddings, represents the th output feature embedding; next, the tracking model inputs this feature vector into a multi-layer perceptron for task classification; the classification tasks include five types of visible light object tracking tasks, visible light-depth object tracking tasks, visible light-infrared object tracking tasks, visible light-event object tracking tasks, and visible light-natural language object tracking tasks; the formula for task classification is:
[0037]
[0038] Among them, represents a multi-layer perceptron through which the multi-modal object tracking model predicts which of the above five tasks the current task belongs to; the output is a task probability distribution, representing the prediction probability of the multi-modal object tracking model for each task; afterwards, the cross-entropy loss function is used to calculate the gap between the prediction result of the multi-modal object tracking model and the true task label, so as to guide the learning of the multi-modal object tracking model; the expression of the cross-entropy loss function is:
[0039]
[0040] Among them, represents the number of tasks, ; is the true label of task , is the probability that the multi-modal object tracking model predicts that the current data belongs to task ; by minimizing the cross-entropy loss function, the multi-modal object tracking model can better distinguish which task and modality the current data belongs to, thereby improving the understanding of the data and further promoting the overall tracking performance;
[0041] Finally, the loss function for training the multi-modal object tracking model is:
[0042]
[0043] Among them, represents the weighted focal loss for foreground / background classification, represents the generalized intersection over union loss, is norm loss, is the cross-entropy loss for task recognition; and are the regularization coefficients of the generalized intersection over union loss and loss respectively;
[0044] Step 5: In the inference stage, a multi-template strategy is adopted; two templates are used: one is a static initial template, and the other is a template that is dynamically updated during the tracking process; the template update mechanism decides when to update based on a fixed time interval and a confidence threshold.
[0045] The above-mentioned public datasets include VastTrack, LaSOT, GOT-10k, TrackingNet, COCO, DepthTrack, Lasher, VisEvent, TNL2K.
[0046] Advantages of the present invention:
[0047] (1) The multi-modal unified representation method proposed by the present invention makes it possible to train an object tracking model that can uniformly process multiple modal signals. The developed multi-modal unified tracking model can handle five different modal tracking tasks through a shared architecture and model, simplifying the development process and model architecture, promoting knowledge sharing, and thus enhancing the flexibility and accuracy of the tracking model.
[0048] (2) The task recognition-assisted training strategy proposed by the present invention promotes the effect of multi-modal task unified training. The proposed soft label type encoding provides the model with more explicit label type information to avoid confusion in label embedding, and ultimately further improves the accuracy of the multi-modal unified tracking model. Brief Description of the Drawings
[0049] Figure 1 It is a framework diagram of a multi-modal tracking method based on modal unified representation. Detailed Embodiment
[0050] The following further describes the detailed embodiments of the present invention in combination with the drawings and technical solutions.
[0051] A multi-modal unified object tracking method based on modal unified representation, the steps are as follows:
[0052] Step 1: Data Preparation
[0053] Collect training data from five multi-modal tracking tasks, including visible light, visible light-depth, visible light-infrared, visible light-event, and visible light-natural language datasets. Generate samples by pairing different modal data, and enhance data diversity by means such as random frame sampling and data augmentation. In addition, appropriately expand the target bounding box to generate template and search area image pairs.
[0054] Step 2: Modal Unified Representation
[0055] Use a multi-modal embedding layer to map input data of different modalities into a unified token embedding representation. The specific operations include:
[0056] (1) Channel concatenation of visible light and other modal data;
[0057] (2) Fix the image patch size, segment and flatten the concatenated data;
[0058] (3) Use a linear transformation to embed the flattened data into a unified representation space;
[0059] (4) For the natural language modality, use the CLIP-L model to extract language features and concatenate them with other modal features to form a unified representation.
[0060] Step 3: Soft Label Type Embedding
[0061] Classify the pixels inside and outside the target bounding box, and introduce soft label weights to adjust the feature embedding. This adjustment mechanism is applicable to both template and search region data, ensuring the accuracy of label embedding, especially in the processing of the foreground and background region boundaries.
[0062] Step 4: Model Training
[0063] Adopt a hybrid training strategy to uniformly train the five-modal task data. Combine the Transformer encoder for feature extraction and fusion, and use the task recognition strategy to enhance the model's ability to distinguish different tasks. At the same time, improve the model performance through the joint optimization of the foreground / background classification loss, bounding box regression loss, and task recognition loss.
[0064] Step 5: Multi-Template Update Strategy in the Inference Stage
[0065] During the inference process, use a combination of static initial templates and dynamically updated templates to continuously adjust the target appearance information and adapt to the appearance changes of the target, thus achieving more stable and accurate target tracking.
Claims
1. A multi-modal unified target tracking method based on modality unified representation, characterized in that: Here are the steps: Step 1: Obtain videos of multimodal target tracking tasks from public datasets for multimodal target tracking model training. The videos of multimodal target tracking tasks contain five modalities: visible light, visible light-depth, visible light-infrared, visible light-event, and visible light-natural language. Among them, the depth, infrared, event, and natural language modalities are paired with the visible light modality, and the depth, infrared, and event modalities are stored in the form of three-channel images. Randomly collect image frames in the video of the multimodal target tracking task, expand the bounding boxes of the targets of interest marked on the image frames by 2 times and 4 times respectively to generate sample pairs of template images and search area images, and use brightness transformation and inversion for augmentation; Step 2: Convert the input signals of different modalities of each multimodal target tracking task into a unified labeled embedding form; (1) For the visible light-depth target tracking task, the visible light-infrared target tracking task, and the visible light-event target tracking task, the depth, infrared, and event modalities are collectively referred to as auxiliary modality data and marked as DTE; the visible light modality data is marked as RGB; After the visible light modality data and the auxiliary modality data are concatenated in the channel direction, a multimodal embedding layer is used to perform a label embedding conversion together, so that the visible light modality data and the auxiliary modality data are converted into a unified label embedding representation; Image I RGB ∈R H×W×3 and image I DTE ∈R H×W×3 Stitching is performed along the channel direction to obtain the stitching image I concat ∈R H ×W×6 , see the following formula: Next, stitch the images I concat The image is divided into fixed-size blocks, each of which is P×P×6; then, each P×P×6 block is flattened into a 6P 2 A one-dimensional vector of; Finally, a linear transformation is applied to map the flattened image block vector to the embedding space, as shown below: E (i) =W p P (i) +b p Among them, E (i) represents the embedding vector of the i-th image block, whose dimension is D; P (i) represents the flattened vector of the i-th image block, W p The dimension is D×6P 2 The weight matrix, b p is a bias term with dimension D; (2) For the object tracking task that does not contain DTE data, a six-channel input is created by copying the three channels of the RGB data, and then processed using the multimodal embedding layer to obtain a multimodal embedding; In the visible light-natural language object tracking task, for the natural language modality, a language model is used as a text encoder to extract a language feature embedding; the language model is a CLIP-L model, on which a linear layer is added to adjust the dimension; the language feature embedding is then concatenated with the multimodal embedding and input into the Transformer encoder; (3) For target tracking tasks that do not contain language modalities, fill in with a fixed sentence; Through the above methods, unified tag embeddings of multimodal search areas, templates, and text descriptions are obtained; these unified tag embeddings are directly sent to a Transformer encoder after being spliced in the spatial dimension; the Transformer encoder completes the joint feature extraction and fusion of these unified tag embeddings through the attention mechanism; Step 3: Add soft marker type encoding to the obtained unified marker embedding, and the foreground marker and background marker of the template image are distinguished by the given bounding box B as follows: Given a template image containing the target and its bounding box B. The template image containing the target is the result of splicing the template image obtained in step 1 in the channel direction of step 2. First, create a mask M∈R with the same size as the template image H×W , in this mask, pixels inside the bounding box are assigned a value of 1, and pixels outside the bounding box are assigned a value of 0: Next, the mask M is divided into non-overlapping image blocks of size P×P; the kth image block is denoted as Then, for each image patch, the average is calculated: in, is the average value of the kth image block; Each marker type corresponds to a learnable marker type embedding, including the foreground marker type embedding E of the template image fg , Template background tag type embed E bg and the tag type embedding E of the search region image search , these tag type embeddings are learned in the training phase and fixed in the inference phase; for the kth image block embedding of the template image, the following adjustments are made: in, represents the adjusted embedding of the kth image patch, E (k) represents the original multimodal image patch embedding; for the search area image, only the search area label type is embedded in E search Added to each image patch embedding, no longer distinguishing between foreground and background: Step 4: Use the above multimodal tag embedding to train a multimodal target tracking model; the multimodal target tracking model consists of a Transformer encoder and a tracking head prediction network; the Transformer encoder is used to extract and fuse the features of the input multimodal search area image and template image, using the HiViT structure; the tracking head prediction network is used to predict the tracking results based on the features output by the Transformer encoder, using the OSTrack head network structure; in order to train the multimodal target tracking model, the following training and optimization strategies are used: Using a multi-task data hybrid training method, we first average all feature embeddings output by the Transformer encoder to generate a single feature vector E avg ; The calculation formula of the eigenvector is: Where N represents the number of output feature embeddings, represents the feature embedding of the i-th output; next, the tracking model inputs the feature vector into a multi-layer perceptron for task classification; the classified tasks include five types of visible light target tracking tasks, visible light-depth target tracking tasks, visible light-infrared target tracking tasks, visible light-event target tracking tasks, and visible light-natural language target tracking tasks; the formula for task classification is: y task =MLP(E avg ) Where MLP(·) represents a multi-layer perceptron; y task is a task probability distribution; then, the cross entropy loss function is used to calculate the gap between the prediction result of the multimodal target tracking model and the true task label; the expression of the cross entropy loss function is: Where K represents the number of tasks, K = 5; is the true label of task j, is the probability that the current data belongs to task j predicted by the multimodal target tracking model; The loss function for multimodal target tracking model training is: in, represents the weighted focal loss for foreground / background classification, represents the generalized intersection-over-union loss, is the l1 norm loss, is the cross entropy loss for task identification; G =2, Step 5: In the inference phase, a multi-template strategy is adopted; two templates are used: one is a static initial template, and the other is a template that is dynamically updated during the tracking process; the template update mechanism decides when to update based on a fixed time interval and confidence threshold.
2. The multi-modal unified target tracking method based on modality unified representation according to claim 1 is characterized in that: Public datasets include VastTrack, LaSOT, GOT-10k, TrackingNet, COCO, DepthTrack, Lasher, VisEvent, and TNL2K.
Citation Information
Patent Citations
Visual learning method based on multi-knowledge fusion
CN115115918A
RGB-T target tracking method based on multi-modal hierarchical relation modeling
CN116580275A
Cited By
A fast-moving small target tracking method based on dual-modal fusion
CN120472377B
Target tracking-oriented adaptive time window event image processing method
CN122416344A
An Adaptive Temporal Window Event Image Processing Method for Target Tracking
CN122416344B