Target tracking method and system based on global semantics and spatial inductive bias
Patent Information
- Application Number
- CN202610813889.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-08
AI Technical Summary
[0004]然而,这种加速策略忽略了特征提取过程中模板语义对搜索分支的引导作用,导致浅层特征中混杂大量背景噪声而缺乏有效的语义提纯;同时,深度剪裁后的纯Transformer架构因缺乏显式的空间定位先验,使得深层自注意力计算在缺乏空间约束的情况下易产生注意力发散
1、通过在特征提取阶段早期嵌入门控全局模板调制模块,以极低计算开销将全局模板语义向量和搜索特征序列进行通道级动态调制,实现了通道级背景噪声的自适应抑制与早期特征的语义提纯,有效缓解了轻量级异步框架中早期语义缺失的问题;
Smart Images

Figure CN122335911B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and image processing technology, and in particular to a target tracking method and system based on global semantics and spatial inductive bias. Background Technology
[0002] Target tracking is one of the fundamental tasks in computer vision, aiming to accurately locate a target in subsequent frames based on a bounding box given in the first frame of a video. With the widespread adoption of drones and mobile robots, target tracking technology is widely used in real-world scenarios such as disaster relief, intelligent transportation, and agricultural inspection. However, targets in real-world environments often face challenges such as drastic deformation, rapid movement, occlusion, changes in lighting, and interference from similar objects. Especially on edge devices, strict power consumption and computing power limitations place extremely high demands on algorithm latency and memory usage. Therefore, designing a lightweight tracking framework that balances high accuracy and high efficiency has become a current research hotspot.
[0003] Early deep tracking methods were primarily based on the Siamese network paradigm of Convolutional Neural Networks (CNNs), tracking by learning the similarity measure between the target and the search region offline. However, CNNs are limited by their local receptive fields, making it difficult to model global contextual relationships, and they are susceptible to background interference in complex scenes, leading to tracking drift. In recent years, visual Transformers, with their powerful global long-range dependency modeling capabilities, have gradually replaced CNNs as the mainstream architecture for target tracking. Single-stream frameworks, such as OSTrack, concatenate templates and search regions into a unified labeled sequence, utilizing self-attention mechanisms for joint deep modeling, significantly improving tracking accuracy. However, the computational complexity of standard self-attention increases quadratically with the number of labels, resulting in massive computational overhead for heavyweight Transformer models, making it difficult to run in real-time on resource-constrained edge devices (typically requiring greater than 30 FPS). To balance accuracy and speed, current research focuses on lightweight target tracking algorithms. Some works employ lightweight backbone networks or introduce neural architecture search techniques to automatically search for efficient tracking structures, effectively reducing the number of model parameters and computational overhead. The LiteTrack approach further proposes layer pruning and asynchronous feature extraction strategies: by performing top-down pruning on heavy Transformers, template and search region are independently computed and cached in shallow layers, which significantly improves the inference frame rate on edge devices.
[0004] However, this acceleration strategy ignores the guiding role of template semantics on the search branch during feature extraction, resulting in shallow features being mixed with a large amount of background noise and lacking effective semantic purification. Simultaneously, the pure Transformer architecture after deep pruning lacks explicit spatial localization priors, making deep self-attention computation prone to attention divergence in the absence of spatial constraints. Therefore, it is necessary to address the tracking drift problems caused by semantic deficiencies in the feature extraction stage and spatial blindness in the interaction stage in existing technologies. Summary of the Invention
[0005] In view of the above situation, the main objective of the present invention is to propose a target tracking method and system based on global semantics and spatial inductive bias to solve the above-mentioned technical problems.
[0006] This invention proposes a target tracking method based on global semantics and spatial inductive bias, the method comprising the following steps: Step 1: Construct a tracking model based on the image patch embedding module, Transformer encoder, gated global template modulation module, correlation-aware spatial gating module, and prediction head; Step 2: Use a large-scale dataset to pre-train the parameters in the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module to obtain a pre-trained tracking model; Step 3: Use the image patch embedding module to perform image patch embedding operation on the template image, and use the Transformer encoder to extract features to obtain the deep template label sequence. Then, perform global average pooling operation on the deep template label sequence to obtain the global template semantic vector. Step 4: Use the image patch embedding module to perform image patch embedding operation on the search image, and use the Transformer encoder to extract features to obtain the search feature sequence. Use the gated global template semantic vector to perform channel-level dynamic modulation on the search feature sequence and the global template semantic vector to obtain the output result calculated by the gated global template modulation module. Step 5: Concatenate the output result calculated by the gated global template modulation module with the deep template label sequence along the spatial dimension to form a joint label sequence; use the correlation-aware spatial gating module to perform spatial mask modulation on the search feature segments in the joint label sequence to obtain the spatially gated modulated search features.
[0007] Step 6: Perform multi-head self-attention calculation on the spatially gated modulated search features, and extract feature segments belonging to the corresponding regions of the search features to reconstruct a two-dimensional feature map. Feed the two-dimensional feature map into the prediction head for decoding to obtain the tracking result. Based on the tracking result, construct a loss function to optimize the pre-trained tracking model and obtain the optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained.
[0008] This invention also proposes a target tracking system based on global semantics and spatial inductive bias, the system comprising: Build modules are used for: A tracking model is built based on an image patch embedding module, a Transformer encoder, a gated global template modulation module, a correlation-aware spatial gating module, and a prediction head. Pre-trained modules are used for: The parameters of the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module are pre-trained using a large-scale dataset to obtain a pre-trained tracking model. Extraction module, used for: The template image is embedded using an image patch embedding module, and features are extracted using a Transformer encoder to obtain a deep template label sequence. The deep template label sequence is then subjected to global average pooling to obtain a global template semantic vector. Modulation module, used for: The image patch embedding module is used to perform image patch embedding operation on the search image, and the Transformer encoder is used to extract features to obtain the search feature sequence. The search feature sequence and the global template semantic vector are input into the gated global template modulation module for channel-level dynamic modulation to obtain the output result calculated by the gated global template modulation module. Gating module, used for: The output calculated by the gated global template modulation module is concatenated with the deep template label sequence along the spatial dimension to form a joint label sequence; the search feature fragments in the joint label sequence are spatially masked and modulated using the correlation-aware spatial gating module to obtain the spatially gated modulated search features. The feature interaction and prediction module is used for: Multi-head self-attention computation is performed on the spatially gated modulated search features, and feature fragments belonging to the corresponding regions of the search features are extracted to reconstruct a two-dimensional feature map. The two-dimensional feature map is then fed into the prediction head for decoding to obtain the tracking result. A loss function is constructed based on the tracking result to optimize the pre-trained tracking model, resulting in an optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained.
[0009] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By embedding a gated global template modulation module early in the feature extraction stage, the global template semantic vector and the search feature sequence are dynamically modulated at the channel level with extremely low computational overhead. This achieves adaptive suppression of channel-level background noise and semantic purification of early features, effectively alleviating the problem of early semantic loss in lightweight asynchronous frameworks. 2. After the joint label sequence is formed, the relevance-aware spatial gating module is used to perform spatial mask modulation on the search feature segments in the joint label sequence. The spatial inductive bias is explicitly injected using dot product operation and convolution smoothing, which guides the multi-head self-attention calculation based on the search features after spatial gating modulation. This guides the attention mechanism to accurately lock the target region, overcoming the defect of blind spatial perception of the network after deep pruning. 3. The proposed channel-level dynamic modulation and spatial mask modulation mechanisms significantly improve the robustness of lightweight trackers in complex interference scenarios without increasing floating-point computation, achieving a better balance between accuracy and speed, and possessing good potential for deployment on edge devices.
[0010] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by means of embodiments of the invention. Attached Figure Description
[0011] Figure 1 This is a flowchart of the target tracking method based on global semantics and spatial inductive bias proposed in this invention.
[0012] Figure 2 This is a structural diagram of the target tracking framework based on global semantics and spatial inductive bias proposed in this invention.
[0013] Figure 3 This is a structural diagram of the gated global template modulation module proposed in this invention.
[0014] Figure 4 This is a structural diagram of the correlation-aware spatial gating module proposed in this invention.
[0015] Figure 5 This is a schematic diagram of the lightweight target tracking system based on global semantics and spatial inductive bias proposed in this invention. Detailed Implementation
[0016] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0017] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to provide some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0018] Please see Figure 1 This invention proposes a target tracking method and system based on global semantics and spatial inductive bias. The method includes the following steps: Step 1: Construct a tracking model based on the image patch embedding module, Transformer encoder, gated global template modulation module, correlation-aware spatial gating module, and prediction head.
[0019] Step 2: Use a large-scale dataset to pre-train the parameters in the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module to obtain a pre-trained tracking model; It should be noted that the parameters in the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module need to be initialized. After initialization, the parameters in the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module are pre-trained using a large-scale dataset to obtain a pre-trained tracking model.
[0020] Step 3: Use the image patch embedding module to perform image patch embedding operation on the template image, and use the Transformer encoder to extract features to obtain the deep template label sequence. Perform global average pooling operation on the deep template label sequence to obtain the global template semantic vector.
[0021] Please see Figure 2 In step 3, the template image is embedded using the image patch embedding module, and features are extracted using the Transformer encoder to obtain a deep template label sequence. Global average pooling is then performed on the deep template label sequence to obtain a global template semantic vector. Specifically, the steps include the following: Perform image block embedding operation on the template image to cut the image into a preset number of non-overlapping template image blocks; The non-overlapping image blocks of the template image are flattened into one-dimensional vectors. The one-dimensional vectors are then projected into a high-dimensional feature space through a linear transformation layer in the image block embedding module, and converted into fixed-dimensional embedding vectors of the template image. Positional encoding is added to the fixed-dimensional embedding vector of the template image to obtain the feature representation of the template image; The template image feature representation is input into the Transformer encoder for feature extraction to obtain a deep template label sequence; Traverse each feature channel in the deep template label sequence, accumulate the feature values corresponding to the spatial location of each channel and take the arithmetic mean to obtain the scalar feature values of all channels; The scalar feature values of all channels are concatenated to obtain the global template semantic vector; In the step of inputting the template image feature representation into the Transformer encoder for feature extraction to obtain the deep template label sequence, the corresponding relationship in the process is as follows: ; in, Represents a deep template marker sequence. Indicates the number of spatial markers, Indicates the total number of channels. Indicates the channel index. Indicating the first element in the characteristic sequence The first spatial marker Characteristic response values of each channel; In the step of traversing each feature channel in the deep template labeling sequence, summing the feature values of all spatial locations under the target channel and taking the arithmetic mean to obtain the scalar feature value, the corresponding relationship in the process is as follows: ; in, Indicates the first Each channel scalar characteristic value; In the step of concatenating the scalar feature values obtained from all channels to obtain the global template semantic vector, the corresponding relationship in the process is as follows: ; in, Represents the global template semantic vector. This represents the scalar characteristic value of the first channel. This represents the scalar characteristic value of the second channel. Indicates the total number of channels. Scalar eigenvalues.
[0022] It should be noted that the obtained deep template tag sequence and global template semantic vector are cached offline, and referenced... Figure 2In the feature extraction (FE) stage, the template image branch includes a 9-layer cascaded Transformer encoder; in the search image branch, the first 3 layers are ordinary Transformer encoders, and the last 3 layers are combined layers consisting of a gated global template modulation module and a Transformer encoder cascaded in sequence, with the combined layer repeated 3 times; in the asynchronous interaction (AI) stage, the combined layer consists of a correlation-aware spatial gating module and a Transformer encoder cascaded in sequence, with the combined layer repeated 3 times.
[0023] In this embodiment, through global average pooling, the global template semantic vector integrates global information about the target and its surrounding environment, which can be used as a prior for subsequent modulation. This vector essentially encodes a highly condensed background distribution prior.
[0024] Step 4: Use the image patch embedding module to perform image patch embedding operation on the search image, and use the Transformer encoder to extract features to obtain the search feature sequence. Use the gated global template semantic vector to perform channel-level dynamic modulation on the search feature sequence and the global template semantic vector to obtain the output result calculated by the gated global template modulation module. Please see Figure 2 and Figure 3 In step 4, the image patch embedding module is used to perform image patch embedding on the search image, and the Transformer encoder is used to extract features to obtain the search feature sequence. The search feature sequence and the global template semantic vector are then subjected to channel-level dynamic modulation using the gated global template modulation module to obtain the output result calculated by the gated global template modulation module. Specifically, the steps are as follows: Perform image block embedding on the search image to cut the image into a preset number of non-overlapping image blocks; The non-overlapping image blocks of the search image are flattened into one-dimensional vectors. The one-dimensional vectors are then projected into a high-dimensional feature space through a linear transformation layer in the image block embedding module, and converted into fixed-dimensional embedding vectors of the search image. Add positional encoding to the fixed-dimensional embedding vector of the search image to obtain the feature representation of the search image; It should be noted that the search image feature representation includes spatial location information; The feature representation of the search image is input into the Transformer encoder for feature extraction to obtain the search feature sequence; The global template semantic vector is input into the gated global template modulation module and subjected to layer normalization to obtain a normalized vector. The normalized vector is then input into the multilayer perceptron of the gated global template modulation module and subjected to nonlinear transformation using the SiLU activation function to obtain an intermediate feature vector. The intermediate feature vector is then divided into a feature scaling vector and a feature translation vector along the channel dimension. The multilayer perceptron includes fully connected layers. A learnable gating scalar parameter is introduced. The learnable gating scalar parameter is multiplied by the feature scaling vector to obtain the first calculation result. The first calculation result is multiplied element-wise by the search feature sequence to obtain the second calculation result. The learnable gating scalar parameter is multiplied by the feature translation vector to obtain the third calculation result. The second calculation result and the third calculation result are added to obtain the output result after calculation by the gating global template modulation module. The process involves inputting the global template semantic vector into the gated global template modulation module, performing layer normalization to obtain a normalized vector, inputting the normalized vector into the multilayer perceptron of the gated global template modulation module, and performing a nonlinear transformation using the SiLU activation function to obtain an intermediate feature vector. The intermediate feature vector is then divided into a feature scaling vector and a feature translation vector along the channel dimension. The multilayer perceptron includes fully connected layers, and the corresponding relationships in this process are as follows: ; in, Represents the feature scaling vector. Represents the feature translation vector. Presentation layer normalization operation, Represents a multilayer perceptron; It should be noted that this multilayer perceptron contains two fully connected layers, with a SiLU activation function used for nonlinear transformation in between. The steps of introducing a learnable gating scalar parameter, multiplying the learnable gating scalar parameter with the feature scaling vector to obtain the first calculation result, multiplying the first calculation result element-wise with the search feature sequence to obtain the second calculation result, multiplying the learnable gating scalar parameter with the feature translation vector to obtain the third calculation result, and adding the second and third calculation results to obtain the output result calculated by the gating global template modulation module, have the following relationships: ; in, This represents the output result calculated by the gated global template modulation module. Indicates the search feature sequence. This indicates an element-wise multiplication operation. This represents a learnable gated scalar parameter.
[0025] It should be noted that the multilayer perceptron maps the global template semantic vector to an intermediate feature vector with doubled dimensions. This is then divided into two modulation vectors along the channel dimension: a feature scaling vector and a feature translation vector; this is done during the initial training phase. This module is equivalent to an identity mapping, preserving the pre-trained weights. As the training process progresses, Through backpropagation and adaptive updates, the network gradually learns to use global template semantic vectors to differentially modulate each channel of the search features.
[0026] Statistical results show that after sufficient training, most feature channels were assigned scaling factors less than 1.0, exhibiting a significant collective suppression trend, with strong selective differences in suppression magnitude between 0.4 and 0.95. This indicates that the gated global template modulation module essentially acts as an adaptive noise reduction filter, effectively suppressing the response of redundant background channels through wide-range, differentiated channel attenuation, thus achieving sparsity purification of early features. The purified features not only effectively reduce the interference of background similarities on subsequent attention calculations but also lay a solid foundation for spatial focusing of the deep correlation perception spatial gating module. As training progresses, Adaptive updates drive the network to suppress redundant channels that respond to background noise, thereby achieving sparsity purification of early features.
[0027] Step 5: Concatenate the output calculated by the gated global template modulation module with the deep template label sequence along the spatial dimension to form a joint label sequence; use the correlation-aware spatial gating module to perform spatial mask modulation on the search feature segments in the joint label sequence to obtain the spatially gated modulated search features. Please see Figure 2 and Figure 4 In step 5, the search feature segments in the joint label sequence are spatially masked and modulated using a correlation-aware spatial gating module to obtain the spatially gated modulated search features. This process includes the following steps: The joint labeled sequence is input into the relevance-aware spatial gating module, and two learnable linear projection matrices are constructed and learnable residual injection parameters are introduced. The two learnable linear projection matrices are used to perform dimensionality reduction on the search feature fragments and global template semantic vectors in the joint labeled sequence to generate query vectors and key vectors. Perform a dot product operation on the query vector and the key vector, and reshape it into a two-dimensional feature map; A convolution operation is performed on the two-dimensional spatial feature map to obtain a smoothed two-dimensional spatial feature map. The sigmoid activation function in the correlation-aware spatial gating module is used to normalize the smoothed two-dimensional spatial feature map and flatten it to obtain a one-dimensional spatial activation mask. The fourth calculation result is obtained by multiplying the one-dimensional spatial activation mask element-wise with the search feature segments in the joint sequence; the fourth calculation result is multiplied with the learnable residual injection parameters to obtain the fifth calculation result; the fifth calculation result is added with the search feature segments in the joint sequence to obtain the spatially gated modulated search features. It should be noted that the two learnable linear projection matrices are two parameters that are gradually learned and fixed after initialization and during the training process in step 2. In the initial stage of building the tracking model in step 1, the computer initializes two random floating-point matrices of dimension C in memory as the initial projection matrices. During the pre-training process in step 2, these two random matrices are multiplied by the input feature segments in forward propagation, and the internal elements of these two matrices are calculated and iteratively updated in round by round using the backpropagation algorithm based on the overall loss function of target tracking until the model training converges. The optimized matrix parameters are then extracted and fixed, and used as fixed linear projection matrices for feature dimensionality reduction calculation in the inference stage of actual target tracking.
[0028] The learnable residual injection parameter is introduced in step 1, in which the present invention strictly initializes the learnable residual injection parameter to 0; in step 2, during the pre-training of the tracking model, this parameter participates in the forward propagation calculation, and through the backpropagation algorithm, based on the overall loss function of target tracking, the residual injection parameter is updated. When the model converges, this parameter will become an optimal static constant.
[0029] In the steps of inputting the joint labeled sequence into the relevance-aware spatial gating module, constructing two learnable linear projection matrices and introducing learnable residual injection parameters, and using the two learnable linear projection matrices to perform dimensionality reduction on the search feature fragments and global template semantic vectors in the joint labeled sequence to generate query vectors and key vectors, the corresponding relationships in the process are as follows: , ; in, Represents the query vector. Represents the key vector. This represents the search feature fragment in the joint labeled sequence. , These represent two learnable linear projection matrices; In the steps of performing a dot product operation on the query vector and the key vector, and reshaping it into a two-dimensional feature map, the corresponding relationship in the process is as follows: ; in, Represents a two-dimensional spatial feature map. This indicates a reshaping operation. This represents the transpose of the key vector. express The spatial dimension in which it is located It means that the total number is The height and width of the two-dimensional spatial grid obtained by restoring the one-dimensional marker sequence; In the steps of performing convolution on a two-dimensional spatial feature map to obtain a smoothed two-dimensional spatial feature map, normalizing the smoothed two-dimensional spatial feature map using the Sigmoid activation function in the correlation-aware spatial gating module, and then flattening the map to obtain a one-dimensional spatial activation mask, the corresponding relationships in the process are as follows: ; in, This represents a one-dimensional spatial activation mask. Indicates the flattening operation. express Convolution operation, This represents the Sigmoid activation function; In the steps of multiplying the one-dimensional spatial activation mask element-wise with the search feature segments in the joint sequence to obtain the fourth calculation result; multiplying the fourth calculation result with the learnable residual injection parameters to obtain the fifth calculation result; and adding the fifth calculation result with the search feature segments in the joint sequence to obtain the spatially gated modulated search features, the corresponding relationships in the process are as follows: ; in, This represents the search features after spatial gating modulation. This represents the learnable residual injection parameters.
[0030] It should be noted that the dimensionality reduction operation of the search feature fragments and global template semantic vector in the joint label sequence is performed by using two learnable linear projection matrices, where the feature dimension is compressed to one-quarter of the original dimension to generate query vectors and key vectors, thereby reducing the amount of subsequent floating-point operations while filtering out redundant details. This represents a reshaping operation, used to convert a one-dimensional labeled sequence into a two-dimensional spatial feature map matrix with height and width attributes; since the dot product similarity map is prone to generating high-frequency isolated noise points due to cluttered backgrounds, a... The two-dimensional convolutional kernels are used for spatial smoothing and context aggregation. This convolutional operation not only effectively filters out isolated noise, but also explicitly injects local spatial inductive bias into the pure Transformer architecture, compensating for the shortcomings of the self-attention mechanism in local structure modeling; in the initial training stage This module is equivalent to an identity mapping, ensuring the stability of the pre-trained weights.
[0031] As training progresses, Adaptive growth and spatial activation masks are smoothly injected into search feature fragments in the joint sequence via residual connections. In challenging scenarios, this effectively suppresses responses from cluttered background regions, exhibiting high activation values in the core region containing the target. After modulation, multi-head self-attention computation is performed based on the spatially gated modulated search features. This approach injects an explicit spatial focusing prior into deep self-attention with extremely low parameters and computational cost, guiding the attention mechanism to accurately lock onto the target region.
[0032] Step 6: Perform multi-head self-attention calculation on the spatially gated modulated search features, and extract feature segments belonging to the corresponding regions of the search features to reconstruct a two-dimensional feature map. Feed the two-dimensional feature map into the prediction head for decoding to obtain the tracking result. Based on the tracking result, construct a loss function to optimize the pre-trained tracking model and obtain the optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained. In step 6, a loss function is constructed based on the tracking results to optimize the pre-trained tracking model, resulting in an optimized tracking model. This includes the following steps: Based on the tracking results, weighted focus loss, generalized intersection-union loss and L1 loss are constructed respectively. The pre-trained tracking model is jointly optimized using weighted focus loss, generalized intersection-union loss, and L1 loss to obtain the optimized tracking model. It should be noted that the head network, or prediction head, performs multi-head self-attention calculation on the search features after spatial gating modulation, and extracts feature fragments belonging to the corresponding regions of the search features to reconstruct a two-dimensional feature map. This two-dimensional feature map integrates the discriminative semantic information of the template target with the global contextual information of the search region, and has strong discriminative power and spatial localization accuracy. The fully convolutional prediction head network includes three convolutional branches: the first branch is the center classification branch, which is used to generate a centrality score map, where the value at each position in the score map represents the confidence that the target center is located at that position; the second branch is the offset regression branch, which is used to calculate the discretization error of the target center position to correct the quantization error caused by feature map downsampling; and the third branch is the scale regression branch, which is used to predict the height and width of the target bounding box.
[0033] The loss function is constructed based on the tracking results to optimize the pre-trained tracking model. The specific steps include the following: We construct a weighted focus loss based on the output of the center classification branch, and construct L1 loss and generalized intersection-over-union loss based on the offset regression branch, scale regression branch, and tracking results. There is also a total loss function, and the corresponding relationship in the process is as follows: ; in, Represents the total loss function. Indicates the weighted focus loss. This indicates the generalized intersection and comparison loss. express loss, and These represent the weighting coefficients used in the balanced model optimization. Using weighted focus loss, The pre-trained tracking model is optimized using loss and generalized intersection-union loss to obtain an optimized tracking model, which is then used for target tracking.
[0034] It should be noted that the overall loss function is used during the model training phase to measure the difference between the tracking results output by the tracking model and the actual results, so that the model can output accurate tracking results. Given The loss is sensitive to the target scale. The loss exhibits scale invariance. This invention follows the standard settings in the classic anchorless visual Transformer encoder architecture, setting the weight coefficients for the balanced model optimization to... and This is to ensure the stable convergence of the gradient during backpropagation.
[0035] Furthermore, as a further preferred embodiment of the present invention, the gated global template modulation module is inserted at layers 3, 4, and 5, rather than at shallower layers such as layers 0, 1, and 2, or at all layers. This is because the features extracted at layers 0-2 are mainly low-level texture and edge information, and have not yet formed stable semantic aggregation. Introducing global template priors too early can easily lead to misfiltering of low-level discriminative details; while insertion at all layers would introduce redundant computational overhead and have limited accuracy improvement. At layers 3-5, the low-level feature construction is complete, and the features begin to have preliminary semantics. At this point, introducing global template semantic filtering can achieve the optimal effect of feature purification.
[0036] Furthermore, as a further preferred embodiment of the present invention, the relevance-aware spatial gating module injects at each layer (layers 6, 7, and 8) during the asynchronous interaction phase, rather than intervening only once at the first or last layer. Intervening only at the first layer results in multiple self-attention calculations in subsequent layers gradually diluting the guiding effect of spatial priors; intervening only at the last layer means that attention divergence has already occurred in the preceding layers, making it difficult to recover later. Continuously injecting spatial masks before each layer's attention calculation effectively establishes stable spatial anchor points, continuously suppresses attention divergence, and achieves optimal spatial focusing.
[0037] Please see Figure 5 The present invention also proposes a target tracking system based on global semantics and spatial inductive bias, the system comprising: Build modules are used for: A tracking model is built based on an image patch embedding module, a Transformer encoder, a gated global template modulation module, a correlation-aware spatial gating module, and a prediction head. Pre-trained modules are used for: The parameters of the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module are pre-trained using a large-scale dataset to obtain a pre-trained tracking model. Extraction module, used for: The template image is embedded using an image patch embedding module, and features are extracted using a Transformer encoder to obtain a deep template label sequence. The deep template label sequence is then subjected to global average pooling to obtain a global template semantic vector. Modulation module, used for: The image patch embedding module is used to perform image patch embedding operation on the search image, and the Transformer encoder is used to extract features to obtain the search feature sequence. The search feature sequence and the global template semantic vector are input into the gated global template modulation module for channel-level dynamic modulation to obtain the output result calculated by the gated global template modulation module. Gating module, used for: The output calculated by the gated global template modulation module is concatenated with the deep template label sequence along the spatial dimension to form a joint label sequence; the search feature fragments in the joint label sequence are spatially masked and modulated using the correlation-aware spatial gating module to obtain the spatially gated modulated search features. The feature interaction and prediction module is used for: Multi-head self-attention computation is performed on the spatially gated modulated search features, and feature fragments belonging to the corresponding regions of the search features are extracted to reconstruct a two-dimensional feature map. The two-dimensional feature map is then fed into the prediction head for decoding to obtain the tracking result. A loss function is constructed based on the tracking result to optimize the pre-trained tracking model, resulting in an optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained.
[0038] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0039] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0040] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A target tracking method based on global semantics and spatial inductive bias, characterized in that, The method includes the following steps: Step 1: Construct a tracking model based on the image patch embedding module, Transformer encoder, gated global template modulation module, correlation-aware spatial gating module, and prediction head; Step 2: Use a large-scale dataset to pre-train the parameters in the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module to obtain a pre-trained tracking model; Step 3: Use the image patch embedding module to perform image patch embedding operation on the template image, and use the Transformer encoder to extract features to obtain the deep template label sequence. Then, perform global average pooling operation on the deep template label sequence to obtain the global template semantic vector. Step 4: Perform image patch embedding on the search image using the image patch embedding module, and extract features using the Transformer encoder to obtain the search feature sequence. Then, use the gated global template modulation module to perform channel-level dynamic modulation on the search feature sequence and the global template semantic vector to obtain the output result calculated by the gated global template modulation module. Specifically, this includes: inputting the global template semantic vector into the gated global template modulation module and performing layer normalization to obtain a normalized vector; inputting the normalized vector into the multilayer perceptron of the gated global template modulation module and performing nonlinear transformation using the SiLU activation function to obtain an intermediate feature vector; and dividing the intermediate feature vector equally along the channel dimension into a feature scaling vector and a feature translation vector. A learnable gating scalar parameter is introduced. The learnable gating scalar parameter is multiplied by the feature scaling vector to obtain the first calculation result. The first calculation result is multiplied element-wise by the search feature sequence to obtain the second calculation result. The learnable gating scalar parameter is multiplied by the feature translation vector to obtain the third calculation result. The second calculation result and the third calculation result are added to obtain the output result after calculation by the gating global template modulation module. Step 5: Concatenate the output result calculated by the gated global template modulation module with the deep template label sequence along the spatial dimension to form a joint label sequence; use the correlation-aware spatial gating module to perform spatial mask modulation on the search feature segments in the joint label sequence to obtain the spatially gated modulated search features; Step 6: Perform multi-head self-attention calculation on the spatially gated modulated search features, and extract feature segments belonging to the corresponding regions of the search features to reconstruct a two-dimensional feature map. Feed the two-dimensional feature map into the prediction head for decoding to obtain the tracking result. Based on the tracking result, construct a loss function to optimize the pre-trained tracking model and obtain the optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained.
2. The target tracking method based on global semantics and spatial inductive bias according to claim 1, characterized in that, In step 3, the template image is embedded using an image patch embedding module, and features are extracted using a Transformer encoder to obtain a deep template label sequence. Global average pooling is then performed on the deep template label sequence to obtain a global template semantic vector. Specifically, the steps include: Perform image block embedding operation on the template image to cut the image into a preset number of non-overlapping template image blocks; The non-overlapping image blocks of the template image are flattened into one-dimensional vectors. The one-dimensional vectors are then projected into a high-dimensional feature space through a linear transformation layer in the image block embedding module, and converted into fixed-dimensional embedding vectors of the template image. Positional encoding is added to the fixed-dimensional embedding vector of the template image to obtain the feature representation of the template image; The template image feature representation is input into the Transformer encoder for feature extraction to obtain a deep template label sequence; Traverse each feature channel in the deep template label sequence, accumulate the feature values corresponding to the spatial location of each channel and take the arithmetic mean to obtain the scalar feature values of all channels; By concatenating the scalar feature values of all channels, a global template semantic vector is obtained.
3. The target tracking method based on global semantics and spatial inductive bias according to claim 2, characterized in that, The template image feature representation is input into the Transformer encoder for feature extraction to obtain a deep template label sequence. The corresponding relationship in the process is as follows: ; in, Represents a deep template marker sequence. Indicates the number of spatial markers, Indicates the total number of channels. Indicates the channel index. Indicating the first element in the characteristic sequence The first spatial marker Characteristic response values of each channel; By traversing each feature channel in the deep template labeling sequence, summing the feature values of all spatial locations under the target channel and taking the arithmetic mean, the scalar feature value is obtained. The corresponding relationship in this process is as follows: ; in, Indicates the first Each channel scalar characteristic value; By concatenating the scalar feature values obtained from all channels, a global template semantic vector is obtained. The corresponding relationship in this process is as follows: ; in, Represents the global template semantic vector. This represents the scalar characteristic value of the first channel. This represents the scalar characteristic value of the second channel. Indicates the total number of channels. Scalar eigenvalues.
4. The target tracking method based on global semantics and spatial inductive bias according to claim 3, characterized in that, In step 4, the image patch embedding module is used to perform image patch embedding on the search image, and the Transformer encoder is used to extract features to obtain the search feature sequence. The search feature sequence and the global template semantic vector are then subjected to channel-level dynamic modulation using the gated global template modulation module to obtain the output result calculated by the gated global template modulation module. Specifically, the steps are as follows: Perform image block embedding on the search image to cut the image into a preset number of non-overlapping image blocks; The non-overlapping image blocks of the search image are flattened into one-dimensional vectors. The one-dimensional vectors are then projected into a high-dimensional feature space through a linear transformation layer in the image block embedding module, and converted into fixed-dimensional embedding vectors of the search image. Add positional encoding to the fixed-dimensional embedding vector of the search image to obtain the feature representation of the search image; The search image feature representation is input into the Transformer encoder for feature extraction to obtain the search feature sequence.
5. The target tracking method based on global semantics and spatial inductive bias according to claim 4, characterized in that, The global template semantic vector is input into the gated global template modulation module and subjected to layer normalization to obtain a normalized vector. This normalized vector is then input into the multilayer perceptron of the gated global template modulation module, where a SiLU activation function is used for nonlinear transformation to obtain an intermediate feature vector. This intermediate feature vector is then divided along the channel dimension into a feature scaling vector and a feature translation vector. The corresponding relationship in this process is as follows: ; in, Represents the feature scaling vector. Represents the feature translation vector. Presentation layer normalization operation, Represents a multilayer perceptron; A learnable gating scalar parameter is introduced. This parameter is multiplied by the feature scaling vector to obtain the first result. The first result is then multiplied element-wise by the search feature sequence to obtain the second result. The learnable gating scalar parameter is then multiplied by the feature translation vector to obtain the third result. Finally, the second and third results are added to obtain the output after processing by the gating global template modulation module. The relationships in this process are as follows: ; in, This represents the output result calculated by the gated global template modulation module. Indicates the search feature sequence. This indicates an element-wise multiplication operation. This represents a learnable gated scalar parameter.
6. The target tracking method based on global semantics and spatial inductive bias according to claim 5, characterized in that, In step 5, the search feature segments in the joint label sequence are spatially masked and modulated using a correlation-aware spatial gating module to obtain spatially gated modulated search features. This specifically includes the following steps: The joint labeled sequence is input into the relevance-aware spatial gating module, and two learnable linear projection matrices are constructed and learnable residual injection parameters are introduced. The two learnable linear projection matrices are used to perform dimensionality reduction on the search feature fragments and global template semantic vectors in the joint labeled sequence to generate query vectors and key vectors. Perform a dot product operation on the query vector and the key vector, and reshape it into a two-dimensional feature map; A convolution operation is performed on the two-dimensional spatial feature map to obtain a smoothed two-dimensional spatial feature map. The sigmoid activation function in the correlation-aware spatial gating module is used to normalize the smoothed two-dimensional spatial feature map and flatten it to obtain a one-dimensional spatial activation mask. The fourth calculation result is obtained by multiplying the one-dimensional spatial activation mask element-wise with the search feature fragment in the joint sequence; the fourth calculation result is multiplied with the learnable residual injection parameters to obtain the fifth calculation result; the fifth calculation result is added with the search feature fragment in the joint sequence to obtain the spatially gated modulated search feature.
7. The target tracking method based on global semantics and spatial inductive bias according to claim 6, characterized in that, The joint labeled sequence is input into the relevance-aware spatial gating module, and two learnable linear projection matrices are constructed, along with learnable residual injection parameters. The two learnable linear projection matrices are used to perform dimensionality reduction on the search feature fragments and global template semantic vectors in the joint labeled sequence, generating query vectors and key vectors. The corresponding relationship in this process is as follows: , ; in, Represents the query vector. Represents the key vector. This represents the search feature fragment in the joint labeled sequence. , These represent two learnable linear projection matrices; The query vector and key vector are multiplied by a dot product and then reshaped into a two-dimensional feature map. The corresponding relationship in this process is as follows: ; in, Represents a two-dimensional spatial feature map. This indicates a reshaping operation. This indicates the transpose operation. express The spatial dimension in which it is located It means that the total number is The height and width of the two-dimensional spatial grid obtained by restoring the one-dimensional marker sequence; A convolution operation is performed on the two-dimensional spatial feature map to obtain a smoothed two-dimensional spatial feature map. The smoothed two-dimensional spatial feature map is then normalized using the Sigmoid activation function in the correlation-aware spatial gating module, and flattened to obtain a one-dimensional spatial activation mask. The corresponding relationship in this process is as follows: ; in, This represents a one-dimensional spatial activation mask. Indicates the flattening operation. express Convolution operation, This represents the Sigmoid activation function; The fourth calculation result is obtained by performing element-wise multiplication of the one-dimensional spatial activation mask with the search feature fragments in the joint sequence; the fourth calculation result is then multiplied with the learnable residual injection parameters to obtain the fifth calculation result; finally, the fifth calculation result is added to the search feature fragments in the joint sequence to obtain the spatially gated modulated search features. The corresponding relationship in this process is as follows: ; in, This represents the search features after spatial gating modulation. This represents the learnable residual injection parameters.
8. The target tracking method based on global semantics and spatial inductive bias according to claim 7, characterized in that, In step 6, a loss function is constructed based on the tracking results to optimize the pre-trained tracking model, resulting in an optimized tracking model. This step specifically includes the following steps: Based on the tracking results, weighted focus loss, generalized intersection-union loss and L1 loss are constructed respectively. The pre-trained tracking model is jointly optimized using weighted focus loss, generalized crossover union loss, and L1 loss to obtain the optimized tracking model.
9. A target tracking system based on global semantics and spatial inductive bias, characterized in that, The system employs the target tracking method based on global semantics and spatial inductive bias as described in any one of claims 1-8, and the system comprises: Build modules are used for: A tracking model is built based on an image patch embedding module, a Transformer encoder, a gated global template modulation module, a correlation-aware spatial gating module, and a prediction head. Pre-trained modules are used for: The parameters of the Transformer encoder, the gated global template modulation module, and the correlation-aware spatial gating module are pre-trained using a large-scale dataset to obtain a pre-trained tracking model. Extraction module, used for: The template image is embedded using an image patch embedding module, and features are extracted using a Transformer encoder to obtain a deep template label sequence. The deep template label sequence is then subjected to global average pooling to obtain a global template semantic vector. Modulation module, used for: The image patch embedding module is used to perform image patch embedding operation on the search image, and the Transformer encoder is used to extract features to obtain the search feature sequence. The search feature sequence and the global template semantic vector are input into the gated global template modulation module for channel-level dynamic modulation to obtain the output result calculated by the gated global template modulation module. Gating module, used for: The output calculated by the gated global template modulation module is concatenated with the deep template label sequence along the spatial dimension to form a joint label sequence; the search feature fragments in the joint label sequence are spatially masked and modulated using the correlation-aware spatial gating module to obtain the spatially gated modulated search features. The feature interaction and prediction module is used for: Multi-head self-attention computation is performed on the spatially gated modulated search features, and feature fragments belonging to the corresponding regions of the search features are extracted to reconstruct a two-dimensional feature map. The two-dimensional feature map is then fed into the prediction head for decoding to obtain the tracking result. A loss function is constructed based on the tracking result to optimize the pre-trained tracking model, resulting in an optimized tracking model. The optimized tracking model is used to track the target, and the final tracking result is obtained.
Citation Information
Patent Citations
Target tracking method and system based on global and local two-way extraction and asynchronous enhancement
CN120894562A
Method for re-recognizing object image based on multi-feature information capture and correlation analysis
US20220415027A1