A visual target tracking method and device based on dynamic space-time selection

CN122415681BActive Publication Date: 2026-09-29INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610856819.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-29
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

然而,这种策略存在明显的局限性:一方面,初始参考图像无法适应目标的外观变化;另一方面,盲目引入历史帧可能导致低质量信息污染特征库,引发“漂移”现象

Benefits of technology

[0027]本发明提出了一种基于动态时空选择的视觉目标跟踪方法,旨在提高复杂场景下目标跟踪的准确性和鲁棒性。其核心在于通过目标状态分析器和动态选择器,实现对参考图像特征的筛选增强与时序特征融合,并构建综合参考图像特征以指导搜索区域特征的目标定位。具体流程如下:首先从视频序列中提取模板和搜索区域图像,并利用特征提取网络将它们转换为令牌嵌入。随后,利用目标状态分析器主动识别并放大目标的显著性特征,有效解决了特征被背景淹没的问题。接着,通过动态选择器,结合时序位置编码和门控残差机制,动态筛选并聚合高质量的历史时空信息,避免了低质量历史帧对跟踪器的误导。最终,将交互后的搜索区域特征输入预测头,利用分类和回归网络分别判断跟踪器是否正确识别目标并预测目标的包围框。该方法通过显式的特征增强和时序筛选,有效提升了跟踪器应对遮挡、形变及快速运动等复杂场景的跟踪准确性和抗干扰能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415681B_ABST
    Figure CN122415681B_ABST
Patent Text Reader

Abstract

The application discloses a visual target tracking method and device based on dynamic space-time selection, and belongs to the technical field of visual target tracking. The method first uses a target state analyzer to perform deep state analysis on multiple reference images, generates a space weight map reflecting target saliency, and dynamically enhances key features while suppressing background interference through a soft weighting strategy; then, through a dynamic selector, the enhanced multiple historical information is subjected to time sequence self-attention interaction in combination with time sequence position coding, and the correlation score of the reference image is calculated using the search area image features, high-quality space-time context information is dynamically screened and aggregated based on the dual criteria of saliency and correlation, and accurate target feature representation is constructed; finally, the foreground confidence and the bounding box position of the target are calculated based on the representation. The application significantly improves the accuracy and robustness of target tracking in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of visual target tracking technology, specifically relating to a visual target tracking method and apparatus based on dynamic spatiotemporal selection. Background Technology

[0002] In the field of computer vision, achieving continuous localization of targets of interest in videos is a fundamental and core task. This technology plays a crucial role in scenarios such as intelligent surveillance, autonomous driving, and human-computer interaction.

[0003] Despite significant advancements in tracking algorithms in recent years, maintaining robustness in complex real-world scenarios remains extremely challenging, such as when the target is occluded for extended periods, undergoes drastic deformation, or is surrounded by similar obstructions. Traditional Siamese networks or Transformer-based trackers often rely on a single initial reference image or a simple time-sliding window to update the reference image. However, this strategy has clear limitations: firstly, the initial reference image cannot adapt to changes in the target's appearance; secondly, blindly introducing historical frames can lead to low-quality information contaminating the feature library, causing a "drift" phenomenon.

[0004] Therefore, how to effectively filter and enhance key features while making full use of timing information and suppressing background noise and interference has become a bottleneck problem in improving tracker performance. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a visual target tracking method and apparatus based on dynamic spatiotemporal selection. The technical solution adopted by this invention is as follows:

[0006] A visual target tracking method based on dynamic spatiotemporal selection includes:

[0007] Step S110: Construct a reference image set containing the initial frame and dynamic frames, and acquire subsequent frames after the first frame of the video sequence in real time as the search area image;

[0008] Step S120: Construct a visual tracking model comprising a feature extraction network, a spatiotemporal feature fusion network, and a target prediction head; the feature extraction network is used to convert the image into a feature representation; the spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features; the target prediction head is used to identify and locate the target in the search area image after the interaction between the comprehensive reference image features and the search area features;

[0009] Step S120 includes:

[0010] Step S120-1: Initialize the parameters of the visual tracking model and preprocess the input reference image set and the search region image;

[0011] Step S120-2: Use a feature extraction network to extract deep semantic features of the image and generate the corresponding feature map;

[0012] Step S120-3: Input the feature map output by the feature extraction network into the spatiotemporal feature fusion network, and use the proposed target state analyzer to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features;

[0013] Step S120-4: Add temporal coding to the enhanced reference image features, calculate saliency score and relevance score using a dynamic selector, and filter and soft replace the features of each frame based on the saliency score and relevance score. Then, stitch or fuse the filtered features into temporary reference image features. After constructing the temporary reference image features, obtain the comprehensive reference image features through temporal self-attention interaction and gated residual fusion, and perform deep interaction between the comprehensive reference image features and the search region features.

[0014] Step S120-5: Input the interactive search region features into the target prediction head, calculate the foreground confidence and bounding box regression parameters of the target respectively, determine the candidate position of the target in the search region by comparing the foreground confidence, and obtain the bounding box of the target according to the bounding box regression parameters corresponding to the candidate position, thereby realizing the identification and localization of the target in the search region image.

[0015] A visual target tracking device based on dynamic spatiotemporal selection includes:

[0016] Reference image and search region image acquisition module: Constructs a set of reference images containing the initial frame and dynamic frames, and acquires subsequent frames after the first frame of the video sequence in real time as search region images;

[0017] Feature extraction network, spatiotemporal feature fusion network, and target prediction head module: This module constructs a feature extraction network, a spatiotemporal feature fusion network, and a prediction head containing classification and regression networks. The feature extraction network converts the image into a feature representation. The spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features. The target prediction head is used to identify and locate targets in the search area image based on the result of the interaction between the comprehensive reference image features and the search area features.

[0018] The specific process of integrating the features of the reference image and the features of the search region is as follows:

[0019] Initialize the parameters of the visual tracking model and preprocess the input set of reference images and the search region image;

[0020] Deep semantic features of images are extracted using a feature extraction network to generate corresponding feature maps;

[0021] The feature map output by the feature extraction network is input into the spatiotemporal feature fusion network. The proposed target state analyzer is used to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features.

[0022] Temporal coding is added to the reference image features. Features of each frame are filtered and softly replaced based on saliency and relevance scores. The filtered features are then spliced ​​or fused into temporary reference image features. After constructing the temporary reference image features, a comprehensive reference image feature is obtained by temporal self-attention interaction and gated residual fusion. This comprehensive reference image feature is then deeply interacted with the search region features.

[0023] The interactive search region features are input into the target prediction head, and the foreground confidence and bounding box regression parameters of the target are calculated respectively.

[0024] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the visual target tracking method based on dynamic spatiotemporal selection.

[0025] A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the described visual target tracking method based on dynamic spatiotemporal selection.

[0026] The present invention has the following beneficial effects:

[0027] This invention proposes a visual target tracking method based on dynamic spatiotemporal selection, aiming to improve the accuracy and robustness of target tracking in complex scenes. Its core lies in using a target state analyzer and a dynamic selector to enhance and fuse reference image features with temporal features, constructing a comprehensive reference image feature set to guide target localization within the search region. The specific process is as follows: First, template and search region images are extracted from the video sequence, and a feature extraction network is used to convert them into token embeddings. Then, the target state analyzer actively identifies and amplifies the salient features of the target, effectively solving the problem of features being obscured by the background. Next, through the dynamic selector, combined with temporal position encoding and a gated residual mechanism, high-quality historical spatiotemporal information is dynamically filtered and aggregated, avoiding the misleading influence of low-quality historical frames on the tracker. Finally, the interactive search region features are input into the prediction head, and classification and regression networks are used to determine whether the tracker correctly identifies the target and predicts its bounding box. This method, through explicit feature enhancement and temporal selection, effectively improves the tracking accuracy and anti-interference capability of the tracker in complex scenes such as occlusion, deformation, and rapid movement. Attached Figure Description

[0028] Figure 1 This is a flowchart of the visual target tracking method based on dynamic spatiotemporal selection according to the present invention.

[0029] Figure 2 This is a schematic diagram of the target state analyzer of the present invention.

[0030] Figure 3 This is an interactive schematic diagram of the dynamic selector of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0032] This invention provides a visual target tracking method based on dynamic spatiotemporal selection. The overall process of this method is as follows: After acquiring a reference image and a search region image from a video sequence, the image is first converted into a token embedding using a feature extraction network. Then, the features of the reference image are enhanced using a proposed target state analyzer, and fused using a dynamic selector.

[0033] Figure 1 This is a flowchart of the visual target tracking method based on dynamic spatiotemporal selection according to the present invention. Figure 1 As shown, the specific implementation steps of this method are as follows:

[0034] Step S110: Construct a reference image set containing the initial frame and dynamic frames, and acquire subsequent frames after the first frame of the video sequence in real time as the search area image;

[0035] This step may include: initializing a set of reference images using the first frame of the video sequence, specifically by using the first frame as an initial static reference image and copying it as an initial dynamic reference image, maintaining a dynamic queue containing the most recent historical frames during the tracking process, and using subsequent frames of the video sequence as search region images.

[0036] Step S120: Construct a visual tracking model comprising a feature extraction network, a spatiotemporal feature fusion network, and a target prediction head; the feature extraction network is used to convert the image into a feature representation; the spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features; the target prediction head is used to identify and locate the target in the search area image based on the search area features after deep interaction with the comprehensive reference image features;

[0037] The specific interaction process between the reference image and the search region image in the spatiotemporal feature fusion network is as follows:

[0038] Step S120-1: Initialize the parameters of the visual tracking model and preprocess the input reference image set and the search region image;

[0039] Step S120-2: Use a feature extraction network to extract deep semantic features of the image and generate the corresponding feature map;

[0040] Step S120-3: Input the feature map output by the feature extraction network into the spatiotemporal feature fusion network, and use the proposed target state analyzer to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features;

[0041] Step S120-4: Add temporal coding to the reference image features, use a dynamic selector to calculate the saliency score and relevance score of each frame of reference image features, and filter and soft replace the features of each frame based on the saliency score and relevance score. Then, the filtered features are spliced ​​or fused into temporary reference image features. After constructing the temporary reference image features, the comprehensive reference image features are obtained through temporal self-attention interaction and gated residual fusion. The comprehensive reference image features are then deeply interacted with the search region features.

[0042] Step S120-5: Input the interactive search region features into the target prediction head, and calculate the foreground confidence and bounding box regression parameters of the target. The target prediction head includes a regression network and a classification network, where the classification network outputs the foreground confidence of the target, and the regression network outputs the center coordinate offset and size parameters of the target bounding box. The target prediction head can adopt an anchor-free dense prediction method known in the art that does not require preset anchor boxes. It determines the candidate position of the target in the search region based on the foreground confidence response map, and obtains the bounding box of the target based on the bounding box regression parameters corresponding to the candidate position, thereby realizing the identification and localization of the target.

[0043] In step S120-3, the target state analysis of each frame of the input reference image features using the proposed target state analyzer includes:

[0044] The input reference image features are subjected to deep feature mining using a channel attention module or a hybrid attention module to generate a spatial weight map reflecting the saliency of the target. This process captures inter-channel dependencies using global feature descriptors and extracts salient regions in space through convolution operations. Subsequently, a soft-weighting strategy combined with a learnable scaling factor is used to enhance the original features element-wise, highlighting key target regions and suppressing background noise, thus providing a high-quality feature foundation for subsequent dynamic temporal screening.

[0045] The soft-weighting strategy uses the following formula for feature enhancement:

[0046] ,

[0047] in, To enhance the post-features, For reference image features, For spatial saliency maps, It is a learnable scaling factor constrained by the hyperbolic tangent function (Tanh) to prevent feature values ​​from diverging.

[0048] Step S120-4 may further include:

[0049] Step S120-4-1: Add learnable temporal position coding to the enhanced reference image features of each frame to distinguish the temporal identity of the initial frame, historical frames and the current frame;

[0050] Step S120-4-2: Based on the generated spatial weight map and the image features of the search region, a dynamic selector is used to calculate the comprehensive quality score of each frame of the reference image; the comprehensive quality score integrates the saliency score of the reference image itself and the relevance score of the reference image relative to the search region image; the integration can be implemented by weighted integration, and the integration weight can be a preset weight or a learnable weight; high-quality reference image features are selected according to a preset threshold; in particular, for low-quality features that are not selected, a soft replacement strategy is adopted, and the initial template features are used to fill them in, so as to maintain the stability of the feature distribution and prevent zero-value collapse in subsequent calculations;

[0051] Step S120-4-3: The filtered and soft-replaced reference image features are spliced ​​or fused in the channel dimension or sequence dimension to construct a comprehensive reference image feature containing rich spatiotemporal context;

[0052] Step S120-4-4: Using the temporal self-attention mechanism, the initial reference image features are used as queries, and the comprehensive reference image features are used as keys and values ​​to perform temporal context interaction and calculate the fused spatiotemporal context features.

[0053] Step S120-4-5: Using a channel-level gated residual mechanism, the fused spatiotemporal context features are adaptively injected into the initial reference image features to obtain the final comprehensive reference image features, and these features are then deeply interacted with the search region features; wherein, the injection includes: introducing learnable gate parameters corresponding to the feature channel dimensions. The gating parameters and the fused spatiotemporal context features are weighted channel-by-channel and injected into the initial reference image features by adding the residuals to obtain the comprehensive reference image features, which exemplarily satisfies... ,in Features of the initial reference image For the fused spatiotemporal context features, This is element-wise multiplication; the comprehensive reference image features are a feature representation matrix composed of multiple feature vectors, with dimensions of... ,in The length of the token sequence for the reference image. For feature channel dimensions.

[0054] Step S120-4-4 specifically includes: calculating the similarity matrix between the initial reference image features (query) and the comprehensive reference image features (key), obtaining the attention weight after Softmax normalization, and then applying the weight to the comprehensive reference image features (value) for aggregation.

[0055] Step S120-4-5 specifically includes: initializing a zero-valued gating parameter so that at the start of training... The injected term is zero, thus the model primarily relies on highly reliable initial reference image features. Target localization; gating parameters during model training. The gating parameters participate in backpropagation as learnable parameters and are updated in each training iteration based on the gradients obtained from backpropagation. As training iterations proceed, the gating parameters... The value is gradually updated from zero to non-zero, making right The residual injection contribution gradually increases, thereby achieving the gradual introduction of the fused spatiotemporal context features; wherein, the gradual introduction is achieved by applying bounded constraints to the gating parameters, for example, letting ,in For learnable parameters, The gating coefficient is a sigmoid function, which allows the gating coefficient to change continuously and smoothly from small to large during training iterations.

[0056] The following specific example will illustrate the above method in detail:

[0057] To facilitate understanding of this embodiment, the symbols used herein are explained as follows: and These represent the height and width of the reference image after cropping and scaling, respectively. and These represent the height and width of the search region image after cropping and scaling, respectively; the feature extraction network may include a Patch Embedding layer, whose kernel size and stride can be set to... That is, the kernel size is And the step size is This divides the image into several tokens; Indicates the length of the reference image token sequence, which exemplarily satisfies , Indicates the length of the token sequence in the search region, exemplarily satisfying... ; Indicates the feature channel dimension (embedding dimension); This represents the spatial weight graph generated by the target state analyzer, where, and It can correspond to the spatial resolution of the features of a reference image, and can be exemplarily taken as follows: and ; For learnable scaling factors in soft-weighted enhancement; These are the learnable parameters for the corresponding frame in temporal position coding; and These represent the significance score and the relevance score, respectively. Indicates by and The overall quality score obtained by fusion Indicates the filtering threshold; , , These represent the query, key, and value in the attention mechanism, respectively. Indicates the features of the initial reference image. This represents the spatiotemporal context features after fusion. This indicates channel-level learnable gating parameters. These are unconstrained learnable variables representing the gating parameters; during the target prediction phase, and It can represent the center coordinate offset. and It can represent the width and height of the target bounding box.

[0058] Step S110: Construct a reference image set and obtain the search region image.

[0059] First, an initial reference image (usually the first frame), dynamically updated reference images (such as the most recently updated historical frames), and subsequent frames after the first frame of the video sequence are obtained from the video sequence as search region images. These images serve as the input to the network.

[0060] Step S120: Construct a visual tracking model

[0061] The constructed model includes a feature extraction network for feature extraction, a spatiotemporal feature fusion network for feature enhancement and fusion, and a task prediction head for target prediction. The specific process is as follows:

[0062] Step S120-1: Model initialization and preprocessing

[0063] Weights are loaded using a pre-trained backbone network, and the cropped reference image and the search region image are used as input. Specifically, the input reference image has a size of [size missing]. For example, = =128, the search area image size is For example, = =256.

[0064] Step S120-2: Deep Feature Extraction

[0065] A feature extraction network is used to extract deep semantic features from the input reference image and the search region image, respectively, resulting in reference image features and search region features. The specific kernel size of the convolutional layers (such as patch embedding) in the feature extraction network is... Step size is The transformed feature dimensions are: The reference image features are represented as follows: The search region features are represented as , and These are the sequence lengths of the reference image and the search region image, respectively. For example, one could take... =16 and =768, and can be derived from and Determine the length of the token sequence.

[0066] Step S120-3: Target State Analyzer

[0067] like Figure 2 As shown, the extracted features are used as input, and a target state analyzer is used to perform target state analysis on the features of each frame of the input reference image, generating enhanced multi-frame reference image features, and generating a spatial weight map to perform soft weighted enhancement on the reference image features:

[0068] The target state analyzer can be constructed using channel attention modules or hybrid attention modules to evaluate the importance of features from both channel and spatial dimensions. The target state analyzer performs deep feature mining on the input reference image features to generate a spatial weight map reflecting the saliency of the target. Furthermore, a soft-weighting strategy is used to enhance the original features element-wise to highlight key target regions and suppress background noise. This soft-weighting strategy satisfies the following:

[0069] ,

[0070] in, To enhance the post-features, For reference image features, For spatial weighting, It is a learnable scaling factor constrained by the hyperbolic tangent function, used to prevent feature values ​​from diverging.

[0071] Step S120-4: Dynamic Selector

[0072] like Figure 3 As shown, based on the enhanced reference image features and the search region image features, a dynamic selector is used to filter out high-quality reference image features, and spatiotemporal fusion of features is achieved:

[0073] Step S120-4-1: Add learnable temporal position encoding to the enhanced reference image features of each frame. This is to distinguish the temporal identity of the initial frame, historical frames, and the current frame.

[0074] Step S120-4-2: Based on the spatial weight map and the image features of the search region, calculate the comprehensive quality score of each frame of reference image using a dynamic selector. Specifically, the score incorporates the saliency score of the reference image itself. Relevance score of the reference image relative to the search region image The fusion can be achieved using a weighted fusion method, where the fusion weights can be preset weights or learnable weights. Based on a preset threshold... High-quality reference image features are selected. In particular, a soft replacement strategy is used for low-quality features that are not selected to maintain the stability of the feature distribution and prevent zero-value collapse in subsequent calculations.

[0075] Step S120-4-3: The high-quality reference image features that have been filtered and soft-replaced are spliced ​​or fused in the channel dimension or sequence dimension to construct temporary reference image features containing rich spatiotemporal context.

[0076] Step S120-4-4: Utilize a temporal self-attention mechanism, using the features of the initial reference image as the query. Using temporary reference image features as the key Sum Temporal context interaction is performed to calculate the fused spatiotemporal context features. The specific calculations include: calculating the features of the initial reference image (query). ) and temporary reference image features (key) The similarity matrix between the two images is normalized using Softmax to obtain attention weights, which are then applied to the features (values) of the temporary reference image. Aggregation is performed on ).

[0077] Step S120-4-5: Utilize the channel-level gated residual mechanism to fuse the spatiotemporal context features. Adaptively inject features into the initial reference image In this process, comprehensive reference image features are obtained. The gating parameters... It can be initialized to zero, so that the injected terms are zero at the start of training, thus the model mainly relies on the features of the initial reference image. Target localization; gating parameters during model training. The gating parameters participate in backpropagation as learnable parameters and are updated according to the gradient in each training iteration, thus enabling the gating parameters to be used as learnable parameters. The values ​​are gradually updated from zero to non-zero, thereby gradually increasing the residual injection contribution of the injected terms and achieving the desired effect on the fused spatiotemporal context features. The gradual introduction of [this mechanism] allows gating parameters to utilize unconstrained learnable variables. Bounded constraints are applied to achieve smooth changes. The final synthesized reference image features are a feature representation matrix composed of multiple feature vectors, with dimensions of... Form, in which The length of the token sequence for the reference image. For feature channel dimensions.

[0078] Step S120-5: Target Prediction

[0079] The task prediction head receives the search region features after deep interaction with the integrated reference image features as input. The task prediction head can employ an anchor-free dense prediction method known in the art, which does not require pre-set anchor boxes, to output the foreground confidence and bounding box regression parameters of the target. The classification branch outputs the foreground confidence at each location in the search region feature map and forms a confidence response map. Candidate locations of the target in the search region are determined by comparing the foreground confidence scores. The regression branch outputs bounding box regression parameters at the candidate locations and recovers the center coordinate offset of the target's bounding box by combining the spatial coordinate mapping relationship of the candidate locations. and and width and height dimensions and This allows us to obtain the final bounding box of the target in the search area image and achieve target recognition and localization.

[0080] Through the above steps, the present invention effectively solves the problems of inaccurate feature extraction and blind temporal fusion in traditional methods in complex scenarios, and significantly improves tracking performance.

[0081] The present invention also provides a visual target tracking device based on dynamic spatiotemporal selection, comprising:

[0082] Reference image and search region image acquisition module: Constructs a set of reference images containing the initial frame and dynamic frames, and acquires subsequent frames after the first frame of the video sequence in real time as search region images;

[0083] Feature extraction network, spatiotemporal feature fusion network, and target prediction head module: This module constructs a feature extraction network, a spatiotemporal feature fusion network, and a prediction head containing classification and regression networks. The feature extraction network converts the image into a feature representation. The spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features. The target prediction head is used to identify and locate targets in the search area image based on the result of the interaction between the comprehensive reference image features and the search area features.

[0084] The specific interaction process between the reference image and the search region image in the spatiotemporal feature fusion network is as follows:

[0085] Initialize the parameters of the visual tracking model and preprocess the input set of reference images and the search region image;

[0086] Deep semantic features of images are extracted using a feature extraction network to generate corresponding feature maps;

[0087] The feature map output by the feature extraction network is input into the spatiotemporal feature fusion network. The proposed target state analyzer is used to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features.

[0088] Temporal coding is added to the reference image features. Features of each frame are filtered and softly replaced based on saliency and relevance scores. The filtered features are then spliced ​​or fused into temporary reference image features. After constructing the temporary reference image features, a comprehensive reference image feature is obtained through temporal self-attention interaction and gated residual fusion. This feature is then deeply interacted with the search region features.

[0089] The interactively generated search region features are input into the target prediction head, which calculates the target's foreground confidence and bounding box regression parameters. The prediction head includes a regression network and a classification network. The regression network predicts the center coordinate offset and size of the target's bounding box based on the interactively generated search region features, while the classification network determines whether the position output by the visual tracking model is the true target.

[0090] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the visual target tracking method based on dynamic spatiotemporal selection.

[0091] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the visual target tracking method based on dynamic spatiotemporal selection.

[0092] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0093] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0094] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0095] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

Claims

1. A visual target tracking method based on dynamic spatiotemporal selection, characterized in that, include: Step S110: Construct a reference image set containing the initial frame and dynamic frames, and acquire subsequent frames after the first frame of the video sequence in real time as the search area image; Step S120: Construct a visual tracking model that includes a feature extraction network, a spatiotemporal feature fusion network, and a target prediction head; the feature extraction network is used to convert the image into a feature representation; the spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features; The target prediction head is used to identify and locate targets in the search region image based on the result of the interaction between the comprehensive reference image features and the search region features. Step S120 includes: Step S120-1: Initialize the parameters of the visual tracking model and preprocess the input reference image set and the search region image; Step S120-2: Use a feature extraction network to extract deep semantic features of the image and generate the corresponding feature map; Step S120-3: Input the feature map output by the feature extraction network into the spatiotemporal feature fusion network, and use the proposed target state analyzer to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features; Step S120-4: Add temporal coding to the enhanced reference image features, calculate saliency score and relevance score using a dynamic selector, and filter and soft replace the features of each frame based on the saliency score and relevance score. Then, stitch or fuse the filtered features into temporary reference image features. After constructing the temporary reference image features, obtain the comprehensive reference image features through temporal self-attention interaction and gated residual fusion, and perform deep interaction between the comprehensive reference image features and the search region features. Step S120-5: Input the interactive search region features into the target prediction head, calculate the foreground confidence and bounding box regression parameters of the target respectively, determine the candidate position of the target in the search region by comparing the foreground confidence, and obtain the bounding box of the target according to the bounding box regression parameters corresponding to the candidate position, thereby realizing the identification and localization of the target in the search region image.

2. The visual target tracking method based on dynamic spatiotemporal selection according to claim 1, characterized in that, Step S110 includes: initializing a set of reference images using the first frame of the video sequence, specifically including using the first frame as an initial static reference image, copying it as an initial dynamic reference image, maintaining a dynamic queue containing the most recent historical frames during the tracking process, and using subsequent frames of the video sequence as search region images.

3. The visual target tracking method based on dynamic spatiotemporal selection according to claim 1, characterized in that, In step S120-3, the target state analysis of each frame of the input reference image features using the proposed target state analyzer includes: The target state analyzer is used to perform deep feature mining on the input reference image features to generate a spatial weight map reflecting the saliency of the target. The algorithm captures the dependencies between channels using a global feature descriptor and extracts salient regions in space through convolution operations. Subsequently, it enhances the original features element-wise using a soft weighting strategy combined with a learnable scaling factor to highlight key target regions and suppress background noise.

4. The visual target tracking method based on dynamic spatiotemporal selection according to claim 1, characterized in that, The target state analyzer is constructed using a channel attention module or a hybrid attention module to evaluate the importance of features from both channel and spatial dimensions.

5. The visual target tracking method based on dynamic spatiotemporal selection according to claim 3, characterized in that, Step S120-4 specifically includes: Step S120-4-1: Add learnable temporal position coding to the enhanced reference image features of each frame to distinguish the temporal identity of the initial frame, historical frames and the current frame; Step S120-4-2: Based on the generated spatial weight map Based on the features of the search region image, a dynamic selector is used to calculate the overall quality score of each frame of the reference image; the overall quality score integrates the saliency score of the reference image itself and the relevance score of the reference image relative to the search region image, wherein the saliency score is determined by the spatial weight map. The calculation shows that the fusion is achieved using a weighted fusion method, where the fusion weights can be preset weights or learnable weights; high-quality reference image features are selected based on a preset threshold; and low-quality features that are not selected are filled using initial template features. Step S120-4-3: The filtered and soft-replaced reference image features are spliced ​​or fused in the channel dimension or sequence dimension to construct temporary reference image features containing rich spatiotemporal context; Step S120-4-4: Using the temporal self-attention mechanism, the initial reference image features are used as queries, and the temporary reference image features are used as keys and values ​​to perform temporal context interaction and calculate the fused spatiotemporal context features. Step S120-4-5: Using a channel-level gated residual mechanism, the fused spatiotemporal context features are adaptively injected into the initial reference image features to obtain comprehensive reference image features, and these comprehensive reference image features are then deeply interacted with the search region features; wherein, the injection includes: introducing learnable gate parameters corresponding to the feature channel dimensions. The gating parameters and the fused spatiotemporal context features are weighted channel by channel and injected into the initial reference image features by adding the residuals to obtain the comprehensive reference image features. ,in Features of the initial reference image For the fused spatiotemporal context features, This is element-wise multiplication; the comprehensive reference image features are a feature representation matrix composed of multiple feature vectors, with dimensions of... ,in The length of the token sequence for the reference image. For feature channel dimensions.

6. The visual target tracking method based on dynamic spatiotemporal selection according to claim 5, characterized in that, Step S120-4-4 specifically includes: calculating the similarity matrix between the initial reference image features and the comprehensive reference image features, obtaining the attention weights after Softmax normalization, and then applying the weights to the comprehensive reference image features for aggregation; Step S120-4-5 specifically includes: initializing a zero-valued gating parameter so that at the start of training... The injected term is zero, thus the visual tracking model relies on the features of the initial reference image. Target localization; gating parameters during model training. The gating parameters participate in backpropagation as learnable parameters and are updated in each training iteration based on the gradients obtained from backpropagation. As training iterations proceed, the gating parameters... The value is gradually updated from zero to non-zero, making right The contribution of residual injection gradually increases, thereby realizing the gradual introduction of spatiotemporal context features after fusion.

7. The visual target tracking method based on dynamic spatiotemporal selection according to claim 3, characterized in that, The soft-weighting strategy uses the following formula for feature enhancement: , in, To enhance the post-features, For reference image features, For spatial weighting, It is a learnable scaling factor constrained by the hyperbolic tangent function, used to prevent feature values ​​from diverging.

8. The visual target tracking method based on dynamic spatiotemporal selection according to claim 5, characterized in that, The target prediction head includes a regression network and a classification network. The regression network predicts the center coordinate offset and size of the target bounding box based on the features of the search region after interaction, and the classification network determines whether the position output by the visual tracking model is the real target.

9. A visual target tracking device based on dynamic spatiotemporal selection, characterized in that, include: Reference image and search region image acquisition module: Constructs a set of reference images containing the initial frame and dynamic frames, and acquires subsequent frames after the first frame of the video sequence in real time as search region images; Feature extraction network, spatiotemporal feature fusion network, and target prediction head module: This module constructs a feature extraction network, a spatiotemporal feature fusion network, and a prediction head containing classification and regression networks. The feature extraction network converts the image into a feature representation. The spatiotemporal feature fusion network introduces a target state analyzer and a dynamic selector to filter, enhance, and fuse temporal features of the reference image to generate high-quality comprehensive reference image features. The target prediction head is used to identify and locate targets in the search area image based on the result of the interaction between the comprehensive reference image features and the search area features. The specific process of integrating the features of the reference image and the features of the search region is as follows: Initialize the parameters of the visual tracking model and preprocess the input set of reference images and the search region image; Deep semantic features of images are extracted using a feature extraction network to generate corresponding feature maps; The feature map output by the feature extraction network is input into the spatiotemporal feature fusion network. The proposed target state analyzer is used to perform target state analysis on the features of each frame of the input reference image to generate enhanced multi-frame reference image features. Temporal coding is added to the reference image features. Features of each frame are filtered and softly replaced based on saliency and relevance scores. The filtered features are then spliced ​​or fused into temporary reference image features. After constructing the temporary reference image features, a comprehensive reference image feature is obtained by temporal self-attention interaction and gated residual fusion. This comprehensive reference image feature is then deeply interacted with the search region features. The interactive search region features are input into the target prediction head, and the foreground confidence and bounding box regression parameters of the target are calculated respectively.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the visual target tracking method based on dynamic spatiotemporal selection as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the visual target tracking method based on dynamic spatiotemporal selection as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-scale Transform target tracking method based on space-time template updating

    CN117036417A

  • Visual target tracking method and device based on attention alignment

    CN120125615A