Visual target tracking method utilizing time sequence prompt and track guidance
By combining the historical information prompt network with the trajectory regression network, the motion trajectory information of the target is explicitly modeled, which solves the problem of insufficient information fusion in existing visual target tracking methods and improves the accuracy and robustness of target tracking, especially the stability in complex scenes.
Patent Information
- Application Number
- CN202510603265.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing visual target tracking methods find it difficult to effectively fuse historical appearance information and trajectory information in complex scenes, resulting in insufficient target matching accuracy and robustness. Especially when the target appearance changes dramatically, occlusion or background interference increases, the model prediction stability is difficult to guarantee.
A method combining historical information prompt network with trajectory regression network is adopted. High-dimensional backbone features are extracted through Transformer encoder, and Mamba module is used to store historical information. Combined with multi-head attention mechanism and trajectory regression network, the motion trajectory of the target is explicitly modeled to achieve unified information fusion and prediction.
It significantly improves the accuracy and robustness of target tracking, enhances the ability to distinguish between targets and backgrounds, and improves the cross-scene generalization ability in complex scenarios.
Smart Images

Figure CN120672794A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a target tracking method based on the fusion of historical information and trajectory information, in particular to a target tracking method that integrates a historical information prompt network and a trajectory regression network of a Mamba module. Background Art
[0002] Visual object tracking is a key research topic in computer vision. Its core task is to consistently and accurately locate an initially designated target within a continuous video sequence. In recent years, with the rapid development of deep learning techniques, particularly the widespread application of convolutional neural networks (CNNs) and Transformer architectures, visual object tracking methods have made significant progress. However, in complex real-world scenarios, existing methods still face numerous challenges, such as changing target appearance, occlusions, rapid motion, and background clutter, making it difficult to achieve a balanced tracking accuracy and robustness.
[0003] Traditional visual object tracking methods often rely on visual feature matching strategies, such as the Siamese network family of methods (e.g., SiamFC and SiamRPN). These methods achieve target position regression and matching by measuring the similarity of image features between template and search frames. However, these methods generally ignore the inherent temporal nature of video data and rely solely on static image appearance features. These methods are unable to adapt to dynamic scenes such as rapidly changing target appearance, complex background interference, or short-term occlusions, and are prone to matching failures and target drift. On the one hand, most existing methods fail to fully utilize the rich information accumulated by the target in historical frames, resulting in a lack of contextual support for matching in the current frame. This makes it difficult to effectively distinguish the feature differences between the target and the background, significantly reducing matching accuracy, especially in complex environments or under occlusion. On the other hand, traditional methods often ignore the temporal regularity inherent in the target's motion trajectory in the video sequence and fail to effectively model the target's motion trends in previous frames. As a result, when dealing with fast-moving targets, drastically changing scales, or long-term occlusions, the prediction results are prone to offset or jitter, significantly reducing robustness.
[0004] To improve the model's ability to understand the dynamic behavior of targets, some studies have introduced modeling methods based on the Transformer architecture (such as TransT and STARK). These methods, through global attention mechanisms, enhance the ability to model long-range features, thus improving the model's adaptability to complex scenes. However, these methods still have two limitations in target tracking tasks: first, they lack explicit modeling of the target's long-term motion state; second, they lack a unified decoding mechanism that integrates historical visual information with trajectory structure information. Therefore, in practical applications, when the target's appearance changes dramatically, there is significant occlusion, or background interference increases, model prediction stability remains difficult to ensure. Furthermore, most existing methods utilize historical information implicitly, lacking an explicit historical cue mechanism. This inability to fully mine and store keyframe features makes it easy for the model to forget initial target features or misidentify other similar regions in long-sequence tracking. Furthermore, existing methods also lack explicit modeling and infusion of trajectory information. Current methods generally lack designs that explicitly integrate historical trajectories as structured input into the target prediction process, resulting in limited target motion prediction capabilities, especially under conditions of occlusion or sudden motion changes.
[0005] Therefore, how to efficiently and clearly integrate the historical appearance information and temporal trajectory information of the target, fully explore the spatiotemporal continuity in the video sequence, and build a unified and collaborative tracking prediction framework has become an important research direction and a key technical problem that needs to be solved in the current field of visual target tracking.
[0006] To address the above problems, the present invention proposes a target tracking method that integrates a historical information Mamba prompt network and a trajectory regression network. By jointly modeling the static visual representation and dynamic trajectory change trend of the target, the accuracy, robustness and cross-scenario generalization ability of target tracking in complex scenes are significantly improved. Summary of the Invention
[0007] The purpose of the present invention is to overcome the problems existing in the above-mentioned prior art and provide a target tracking method that effectively integrates historical information and trajectory information to improve the matching accuracy and prediction robustness in target tracking tasks.
[0008] To achieve the above objectives, the present invention proposes a visual target tracking method using timing prompts and trajectory guidance, which includes the following core technical solutions:
[0009] A visual target tracking method using temporal cues and trajectory guidance comprises the following steps:
[0010] Step 1: The first frame of the video sequence is used as the template image, and the subsequent frames are sequentially used as search images. Together with the template image, they are input into the Transformer encoder. The Transformer encoder establishes global visual relationships and extracts high-dimensional backbone features containing contextual semantics in the current search image.
[0011] Step 2: When processing the first search image, the high-dimensional backbone features of the current search image are directly input into the Transformer decoder and the history hint Mamba module. When processing the second and subsequent search images, the high-dimensional backbone features of the current search image are respectively input into the history hint decoder module and the history hint Mamba module. The history hint Mamba module combines the spliced feature map obtained from the previous search image to generate the hint value and hint key, and stores them as historical information in the history memory stack.
[0012] Step 3: The predicted bounding box information of the target obtained from the most recent n search frames is combined into historical trajectory information. If the most recent search image is less than n frames, the bounding box information of the target in the template is copied multiple times to make up n frames. The historical trajectory information is then combined with the target position information in the initialized current search frame image and input into the Tranformer decoder. When processing the first search frame, the Tranformer decoder outputs the predicted bounding box of the current frame target based on the high-dimensional backbone features of the current search image, the target position information in the template, and the target position information in the initialized current search frame image.
[0013] When processing the second frame and subsequent search images, the historical hint decoder module retrieves the most relevant hint value and hint key from the historical memory of the historical hint Mamba module, and combines the high-dimensional backbone features of the current frame with the most relevant hint key and hint value in spatiotemporal position encoding, performs a multi-head attention mechanism fusion, obtains the optimized backbone features of the current frame and inputs them into the Tranformer decoder. The Tranformer decoder outputs the predicted bounding box of the current frame target based on the optimized backbone features of the current frame, historical trajectory information, and the target position information in the initialized current frame search image;
[0014] Step 3: Generate a mask of the target based on the predicted bounding box of the current frame target, and splice the mask with the current frame search image to form a spliced feature map of the current frame search image (composed of the RGB image and the mask map of the predicted target) and input it into the history prompt Mamba module until the predicted bounding boxes of the targets of all images in the video sequence are generated.
[0015] Furthermore, the history hint Mamba module is used to extract features of the spliced feature map of the previous search image using Vision Mamba, output features of the same dimension as the backbone features of the previous search image and reshape them into a three-dimensional feature map;
[0016] Then, the spliced feature map of the previous search image is spliced with the three-dimensional feature map along the channel dimension to obtain the fused feature map, which is then fed into multiple residual modules and pooling modules to generate the prompt value.
[0017] At the same time, the high-dimensional backbone features of the previous frame search image are subjected to 1×1 convolution to obtain the prompt key, and the prompt value and prompt key are written into the historical information memory library.
[0018] Furthermore, the processing of the Tranformer decoder includes the following steps:
[0019] Step 2.1: Construct a trajectory regression network, concatenate the historical trajectory information in the unified coordinate system with the initialization position of the current frame to form a trajectory input sequence, map it into a dense vector through the encoder, combine it with the position encoding, and input it into the masked multi-head attention mechanism to output the fused trajectory vector.
[0020] Step 2.2: The fused trajectory vector and the high-dimensional backbone features of the current frame are input into the multi-head attention mechanism again to obtain the image-text fusion target features. The image-text fusion target features are processed by the residual network and the feedforward neural network to generate the final prediction representation of the target;
[0021] Step 2.3: The final prediction of the target is represented by encoding conversion and softmax function, mapping the final feature to the bounding box coordinate information of the target in the current frame.
[0022] Furthermore, the memory library of the history prompt Mamba module has a capacity of 120 entries and is updated using a first-in-first-out (FIFO) strategy.
[0023] Furthermore, the masked multi-head attention mechanism uses an autoregressive masking strategy to ensure that each position prediction depends only on the previous known trajectory.
[0024] Furthermore, in step 2, n is 8.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] (1) This paper proposes a historical information prompt network that uses the Mamba module as a feature extractor to extract fine-grained features of the target in the current frame and stores the extracted target features in a historical information library for use in target matching in subsequent frames. By establishing this historical information library, the target matching process can be provided with more comprehensive and accurate historical information prompts, significantly enhancing the ability to distinguish between targets and backgrounds.
[0027] (2) The present invention proposes a trajectory regression network that fully utilizes the trajectory information accumulated by the target in the previous frame through the trajectory regression decoder, constructs a unified spatiotemporal modeling mechanism, closely combines the dynamic trajectory of the target with the prediction of the current frame, effectively assists the prediction of the current frame bounding box, and greatly improves the accuracy and stability of the prediction.
[0028] The present invention fully mines and utilizes the detailed features of historical frame targets through the historical information prompt network, solves the problem of insufficient contextual information in target matching in the prior art, and improves the ability to distinguish between targets and backgrounds.
[0029] The present invention effectively integrates dynamic trajectory information through a trajectory regression network, so that the current frame prediction has a clearer motion prior, effectively overcoming the defect of insufficient utilization of trajectory information in existing methods and improving the accuracy and robustness of target bounding box prediction.
[0030] In summary, the present invention effectively improves the matching accuracy and prediction robustness in the target tracking process by combining static historical appearance information with dynamic target trajectory information, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Tracking the overall flow chart framework diagram for the present invention;
[0032] Figure 2 This is a schematic diagram of the historical information prompt network innovated by the present invention;
[0033] Figure 3 Schematic diagram of the innovative trajectory regression head network of the present invention. DETAILED DESCRIPTION
[0034] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings so that those skilled in the art can clearly understand and implement the methods described in the present invention.
[0035] like Figure 1 As shown in the figure, this paper proposes a target tracking method that combines a historical information prompt network with a trajectory regression network. This method fully utilizes the temporal nature of video sequences and improves the accuracy of target matching and the robustness of prediction by explicitly modeling the target's appearance information and motion trajectory in historical frames.
[0036] The core of the present invention includes two key modules: (1) a history information prompt network, which consists of a history prompt Mamba module and a history prompt decoder module, and is used to store and extract key historical frame features; and (2) a trajectory regression network (Transformer Decoder), which is used to integrate the motion trajectory of the target and realize position prediction.
[0037] In this embodiment, the first frame of the input video sequence is set as a template image with a size of 192×192×3, which is used to initialize the visual features of the target. Subsequent frames are used as search images with a size of 384×384×3, which are used to locate the continuous position of the target in the video stream.
[0038] Feature extraction stage
[0039] First, the template image and the current search image are input to the Patch Embedding module, which uses a convolution with a kernel size of 16×16 to divide the image into fixed-size image patches and map them into vector sequences. Specifically:
[0040] · The template image is embedded as a vector sequence with a dimension of 144×768;
[0041] · The search image is embedded as a sequence of vectors of dimension 576×768.
[0042] The template features are then concatenated with the search image features and fed as input to a Transformer encoder (ViT) to establish global visual relationships and extract high-dimensional backbone features that incorporate contextual semantics. ViT retains only the output from the search image, resulting in an output feature size of 576×768.
[0043] First frame prediction stage (no historical information)
[0044] For the first frame of the search image, its backbone features are directly input into the Transformer decoder (trajectory regression network). The trajectory information (x1, y1, w1, w2, x1, y1, w1, w2, ..., x1, y1, w1, w2) is initialized using the target position information in the first frame template. This information, combined with the initial target position information (start, x, y, w, h) for the current frame, serves as the initial input to the Transformer decoder (trajectory regression head network). The target trajectory information is stored in a stack of 8 (i.e., 8 sets of target position information, using a first-in, first-out update strategy). After correlation with the backbone features, the predicted bounding box (x2, y2, w2, h2) of the target in the first frame is obtained.
[0045] Next, a target mask is generated based on the predicted box and concatenated with the current frame search image to form a 4-channel fused feature map. This feature map, along with the backbone features, is fed into the History Hint Mamba module to generate a hint value and hint key. This hint value and hint key are then stored as historical information in a history memory stack (size 120, using a first-in, first-out update strategy), providing more comprehensive and accurate historical information hints for subsequent frame target matching.
[0046] Subsequent frame prediction stage (t≥3) (t=1 is the template, t=2 is the first frame search image)
[0047] When processing the subsequent search image for the tth frame, the template image and the current search image are again processed through PatchEmbedding and ViT to extract backbone features (576×768). The current frame backbone features are then input into the historical cue decoder module, which is responsible for retrieving the most relevant cue value and cue key from the historical memory and fusing them with the current features, significantly enhancing the ability to distinguish between target and background features and effectively improving the target representation capability.
[0048] The fused features are fed into the Transformer decoder (trajectory regression network), along with the historical trajectory information (xt-8, yt-8, wt-8, ht-8, …, xt-1, yt-1, wt-1, ht-1) and the initial predicted position of the current frame (start, x, y, w, h) to complete the regression prediction of the target bounding box (xt, yt, wt, ht) of the current frame.
[0049] Subsequently, the mask generated by the prediction box is concatenated with the current frame image to form a feature map with 4 channels, and input into the historical prompt Mamba module together with the backbone features to generate the prompt key-value pairs of the current frame and continue to update the historical prompt information stack.
[0050] Memory mechanism and history prompt management
[0051] During model training and inference, both the historical information prompt and the trajectory information stack adopt a finite-length FIFO strategy (first-in-first-out):
[0052] · The size of the history prompt memory stack is set to 120, with an initial length of 0, and is used to save the prompt key-value pairs of the history frame;
[0053] · The trajectory information stack size is 8, with the template target position repeated 8 times as the initial information, the initial length is 8, and then it is continuously updated with each frame of information.
[0054] As the number of frames increases, old information is automatically eliminated, maintaining the model's optimal perception within a limited historical window. Each frame is predicted using a dual mechanism of historical memory and trajectory assistance, significantly enhancing the continuity, interference resistance, and temporal robustness of the tracking process.
[0055] Model Configuration
[0056] In this method, the backbone feature extractor adopts the Vision Transformer (ViT) architecture and loads ViT-B weights pre-trained on ImageNet to improve the model's expressiveness and convergence speed in visual feature encoding. Furthermore, the Mamba subnetwork in the historically informed Mamba module loads pre-trained weights from the VIM-S model, leveraging its advantages in modeling long sequences. Considering the introduction of a fourth channel feature map generated by the predicted bounding box mask, the input channel count of the first convolutional layer in the Mamba module is increased from the default 3 to 4 to accommodate this input, enabling unified processing and fusion of the RGB image and mask information.
[0057] Training process:
[0058] This method adopts a phased training strategy to ensure that the historical information prompt mechanism and the trajectory assistance mechanism can be gradually and stably established and work together in the target prediction task.
[0059] Phase 1: Historical information prompt target tracking model training (150 rounds)
[0060] In the first phase, the training goal is to build a prediction model with basic matching capabilities while simultaneously training the historical information prompt network (the historical prompt Mamba module and the historical prompt decoder module). The training cycle for this phase is 150 rounds, and the loss function uses the cross entropy and siou function.
[0061] During this phase, the trajectory assistance mechanism is not yet enabled, and there is no target trajectory information. Only target matching and appearance modeling capabilities are established. After training, the model can output highly accurate target bounding boxes, providing a reliable foundation for building and updating the trajectory information stack in subsequent phases, ensuring its effectiveness and usability.
[0062] Phase 2: Target trajectory information-assisted training (30 rounds)
[0063] After completing the first phase of training, the trajectory-assisted mechanism training phase begins. This training cycle lasts for 30 epochs, and the loss function utilizes the cross-entropy and siou functions. During this phase, the Transformer decoder module (the trajectory regression head network) uses historical trajectory information as auxiliary input for training. Trajectory information is maintained using a first-in-first-out (FIFO) strategy, recording the target position information for the last eight consecutive frames. This information is fed into the decoder along with the initial position (start, x, y, w, h) of the current frame to predict the target bounding box for the current frame.
[0064] Decoding strategies and reasoning mechanisms
[0065] Decoding acceleration strategy during the training phase
[0066] To improve training efficiency, the Transformer decoder does not use a step-by-step prediction strategy during training. Traditional methods require predicting each bounding box parameter sequentially (i.e., predicting x first, then y, then w, and finally h), resulting in longer training times.
[0067] To this end, this method introduces a parallel prediction method based on the mask multi-head attention mechanism, which uses known real bounding boxes to construct an input sequence with a dislocation structure:
[0068] Input sequence: (start, x, y, w, h) → Prediction output: (x, y, w, h, end)
[0069] Through the mask-guided self-attention mechanism, the model can learn the joint representation relationship of all coordinates in a single forward propagation, realize parallel training and regression of bounding box parameters, and significantly save training time.
[0070] Step-by-step reasoning strategy for the testing phase
[0071] During the testing phase, to ensure stability and rigor during inference, this method uses a step-by-step inference method (Auto-Regressive Decoding) for bounding box prediction. The specific process is as follows:
[0072] 1. Use the start token to predict the x coordinate of the target box;
[0073] 2. Use (start, x) to predict the y coordinate;
[0074] 3. Use (start, x, y) to predict w;
[0075] 4. Use (start, x, y, w) to predict h;
[0076] 5. Finally output the complete bounding box position (x, y, w, h).
[0077] This method can effectively reduce the accumulation of prediction errors and improve the robustness of target prediction.
[0078] like Figure 2 The historical information prompt network proposed in this paper is shown. It consists of two parts: the historical prompt Mamba module on the left and the historical prompt decoder module on the right. This module is used to extract prompt features of historical moments and store them in memory to assist in feature optimization and matching of the current frame.
[0079] History prompt Mamba module (history prompt generation)
[0080] This module is used to generate the prompt key and prompt value of the T-th frame image:
[0081] 1. The backbone feature output dimension of the T-th frame image is 576×768, where 576=24×24 represents the number of spatial patches and 768 is the channel dimension.
[0082] 2. The backbone features are reshaped into a three-dimensional backbone feature map with a size of (24, 24, 768).
[0083] 3. At the same time, the T frame image is concatenated with the mask map generated by its predicted bounding box to form a 4-channel input image. The Vision Mamba module extracts features and the output dimension is 576×384, which is reshaped to (24,24,384).
[0084] 4. Concatenate the above two feature maps in the channel dimension to form a fused feature map with a size of (24, 24, 1152), that is, 768+384.
[0085] 5. The first residual module has an input channel count of 1152 to accommodate the high-dimensional input of the fused feature map. This module compresses and projects the high-dimensional information to the target channel dimension, and outputs an intermediate fused feature of size (24, 24, 384) through convolution and normalization.
[0086] 6. The intermediate features are input to the pooling module, which performs maximum pooling and average pooling operations on the channel dimension respectively. After fusing the information of the two, the output is still a feature map of (24, 24, 384).
[0087] 7. The output of the pooling module is added to its input through a residual connection and input to the second residual module for further nonlinear transformation and feature refinement. The output feature remains (24, 24, 384).
[0088] 8. Finally, the fused feature map is reshaped into a two-dimensional form to obtain the prompt value (value), with a size of 576×384.
[0089] 9. At the same time, the initial backbone feature map (24, 24, 768) is channel-compressed and reshaped into a two-dimensional form through 1×1 convolution to obtain the prompt key (key), which also has a size of 576×384.
[0090] The prompt key and prompt value are eventually written into the historical prompt memory bank for use in decoding subsequent frames.
[0091] 2. History Hint Decoder Module
[0092] When processing the T+1 frame image, this module reads historical prompt information from the memory library and enhances the current backbone features.
[0093] 1. The input is the backbone features of the T+1th frame image, with a dimension of 576×768.
[0094] 2. First, the feature is input into the linear layer to reduce the number of channels to 384.
[0095] 3. Then, the query feature (query), prompt key (key) and prompt value (value) are weightedly fused with the spatiotemporal position encoding to enhance the temporal perception ability.
[0096] 4. Then, all three are fed into a multi-head attention mechanism to perform an attention-weighted operation and extract the information most relevant to the historical cues.
[0097] 5. The attention output results are sequentially transformed and feature integrated through the residual connection structure and the feed-forward neural network to obtain the optimized target features.
[0098] 6. The final output dimension is kept at 576×384, which is consistent with the original backbone feature size and can be used for subsequent target position regression or further encoding tasks.
[0099] Module Summary
[0100] This module design uses Vision Mamba to collaboratively construct cue values with backbone features. A 1×1 convolution is used to generate cue keys, which are then stored in a historical memory bank. Subsequent frames utilize a multi-head attention mechanism combined with positional encoding to extract key context from historical cues, guiding the optimal learning of target representations. This mechanism not only improves the discriminative ability of target matching but also enhances the model's robustness to temporal variations and occlusion.
[0101] like Figure 3 The structure design and processing flow of the proposed Transformer Decoder trajectory regression head network are presented in this article. This module aims to use historical trajectory information and the current frame to initialize the target position information, and generate the precise bounding box coordinates of the current frame through an autoregressive prediction mechanism.
[0102] 1. Input design and encoding conversion
[0103] The input of this module consists of two parts:
[0104] 1. Trajectory information stack: stores the predicted position information of the target in the previous 8 frames (..., xt-1, yt-1, wt-1, ht-1), a total of 8×4=32 position elements;
[0105] 2. Current frame initialization position information: including the start token and the initialized (x, y, w, h), a total of 5 position elements.
[0106] The above information is concatenated sequentially to form a trajectory position sequence of length 37 (=32+5). Each position element is used as a token, for a total of 37 tokens, forming the input token sequence. This sequence is first mapped into a 768-dimensional dense vector by the encoding conversion module (Embedding Layer), resulting in a trajectory encoding representation with a dimension of (37×768).
[0107] Subsequently, the encoding sequence is added to its corresponding positional encoding to retain the position prior information of each time step in the trajectory sequence, introducing sequence order perception for the subsequent attention mechanism.
[0108] 2. Trajectory Self-Attention Mechanism Processing (Masked Multi-Head Attention)
[0109] The position-encoded trajectory vector sequence is fed into the Masked Multi-head Attention mechanism:
[0110] · The masking mechanism ensures that the prediction process follows an autoregressive strategy, and the current token can only focus on its previous token to avoid information leakage;
[0111] · The attention module at this stage outputs a trajectory prediction vector sequence of length 5 (corresponding to the start, x, y, w, h of the target) with a dimension of (5×768);
[0112] ·The output representation fuses the important features related to the temporal structure in the historical trajectory and the current estimated position information to form a fused trajectory information vector.
[0113] 3. Target Perception and Attention Fusion
[0114] Then, the fused trajectory information vector and the current frame backbone features are input into the multi-head attention mechanism:
[0115] · Backbone features (i.e., current frame image features) provide spatial visual information of the target;
[0116] · Fuse the trajectory information vector as the query (q) and the backbone features as the key-value pairs (k, v) and perform cross attention;
[0117] · The output dimension is (5×768), corresponding to the image-text fusion features of each predicted token.
[0118] 4. Target Position Regression and Prediction
[0119] The above fusion features are then optimized through the following modules in sequence:
[0120] 1. Residual connection module: enhances feature stability;
[0121] 2. Feed-Forward Neural Network: Introduces nonlinear transformation and further feature abstraction;
[0122] 3. Encoding conversion module: maps the 768-dimensional vector to the final coordinate space output dimension 800, consistent with the previous vocabulary;
[0123] 4. Softmax layer: Performs normalization and outputs the probability distribution of each position in a unified coordinate system, which is then restored to the true target bounding box coordinates (x, y, w, h).
[0124] 5. Design Advantages
[0125] This trajectory regression head network design has the following advantages:
[0126] · Autoregressive mask modeling: simulates language to build pattern predictions and avoid future information leakage;
[0127] · Spatiotemporal fusion mechanism: Fusion of visual features and motion trajectories to ensure that bounding box predictions have both structural and dynamic consistency;
[0128] ·Unified decoding architecture: A unified Transformer Decoder simultaneously processes spatial perception and temporal modeling.
[0129] The foregoing is an example of the best mode of carrying out the present invention. Any portion not described in detail herein is common knowledge within the skill of one of ordinary skill in the art. The scope of protection of the present invention is determined by the claims. Any equivalent transformation based on the technical teachings of the present invention is also within the scope of protection of the present invention.
Claims
1. A visual target tracking method using timing prompts and trajectory guidance, characterized in that: The following steps are involved: Step 1: The first frame of the video sequence is used as the template image, and the subsequent frames are sequentially used as search images. Together with the template image, they are input into the Transformer encoder. The Transformer encoder establishes global visual relationships and extracts high-dimensional backbone features containing contextual semantics in the current search image. Step 2: When processing the first search image, the high-dimensional backbone features of the current search image are directly input into the Transformer decoder and the history hint Mamba module. When processing the second and subsequent search images, the high-dimensional backbone features of the current search image are respectively input into the history hint decoder module and the history hint Mamba module. The history hint Mamba module combines the spliced feature map obtained from the previous search image to generate the hint value and hint key, and stores them as historical information in the history memory stack. Step 3: The predicted bounding box information of the target obtained from the latest n frames of search images is combined into historical trajectory information. If the latest search image is less than n frames, the bounding box information of the target in the template is copied multiple times to make up n frames. The historical trajectory information is then combined with the target position information in the initialized current frame search image and input into the Tranformer decoder. When processing the first frame search image, the Tranformer decoder initializes the trajectory information based on the high-dimensional backbone features of the current search image, the position information of the target in the template, and the target position information in the initialized current frame search image to output the predicted bounding box of the current frame target. When processing the second frame and subsequent search images, the historical hint decoder module retrieves the most relevant hint value and hint key from the historical memory of the historical hint Mamba module, and combines the high-dimensional backbone features of the current frame with the most relevant hint key and hint value in spatiotemporal position encoding, performs a multi-head attention mechanism fusion, obtains the optimized backbone features of the current frame and inputs them into the Tranformer decoder. The Tranformer decoder outputs the predicted bounding box of the current frame target based on the optimized backbone features of the current frame, historical trajectory information, and the target position information in the initialized current frame search image; Step 3: Generate a mask of the target based on the predicted bounding box of the current frame target, and splice the mask with the current frame search image to form a spliced feature map of the current frame search image and input it into the history prompt Mamba module until the predicted bounding boxes of the targets of all images in the video sequence are generated.
2. A visual target tracking method using timing prompts and trajectory guidance according to claim 1, characterized in that: The history hint Mamba module is used to extract features of the spliced feature map of the previous search image using Vision Mamba, output features with the same dimension as the backbone features of the previous search image, and reshape them into a three-dimensional feature map; Then, the spliced feature map of the previous search image is spliced with the three-dimensional feature map along the channel dimension to obtain a fused feature map, which is then fed into multiple residual modules and pooling modules to generate a prompt value. At the same time, the high-dimensional backbone features of the previous frame search image are subjected to 1×1 convolution to obtain the prompt key, and the prompt value and prompt key are written into the historical information memory library.
3. The method for visual target tracking using timing prompts and trajectory guidance according to claim 1, characterized in that: The processing of the Tranformer decoder includes the following steps: Step 2.1: Construct a trajectory regression network, concatenate the historical trajectory information in the unified coordinate system with the initialization position of the current frame to form a trajectory input sequence, map it into a dense vector through the encoder, combine it with the position encoding, and input it into the masked multi-head attention mechanism to output the fused trajectory vector. Step 2.2: The fused trajectory vector and the high-dimensional backbone features of the current frame are input into the multi-head attention mechanism again to obtain the image-text fusion target features. The image-text fusion target features are processed by the residual network and the feedforward neural network to generate the final prediction representation of the target; Step 2.3: The final prediction of the target is represented by encoding conversion and softmax function, mapping the final feature to the bounding box coordinate information of the target in the current frame.
4. The method for visual target tracking using timing prompts and trajectory guidance according to claim 1, characterized in that: The memory library of the history prompt Mamba module has a capacity of 120 entries and is updated using a first-in-first-out (FIFO) strategy.
5. The method for visual target tracking using timing prompts and trajectory guidance according to claim 1, characterized in that: The masked multi-head attention mechanism uses an autoregressive masking strategy to ensure that each position prediction depends only on the previous known trajectory.
6. The method for visual target tracking using timing prompts and trajectory guidance according to claim 1, characterized in that: In step 2, n is 8.
Citation Information
Cited By
Window prediction method and system
CN121188224A
Multi-mode visual single-target tracking method and device based on memory prompt, equipment and medium
CN121304734A
Memory prompt-based multi-modal visual single target tracking method, device, equipment and medium
CN121304734B
Target tracking method and system based on trajectory perception
CN121330013A
A target tracking method and system based on trajectory awareness
CN121330013B