Object Tracking Method, Device, Equipment and Medium Based on Pseudo-Multimodal Retrieval

The pseudo multi-modal retrieval method using vector quantization variational autoencoders and dynamic feature selection improves target tracking precision and efficiency in dynamic environments by capturing semantic information and reducing complexity.

CN120147366BActive Publication Date: 2025-07-15SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510629024.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-15
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

The existing target tracking technology has reduced accuracy under high-speed motion, complex background and occlusion, making it difficult to achieve stable tracking, and is susceptible to interference.

Method used

The target tracking method based on pseudo-multimodal retrieval is adopted to construct a codebook through a vector quantization variational autoencoder, feature extraction and progressive dynamic visual feature selection, irrelevant background information, and pseudo-multimodal retrieval and target prediction.

Benefits of technology

It improves the accuracy and efficiency of target tracking, enhances the robustness and adaptability in complex environments, and reduces the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147366B_ABST
    Figure CN120147366B_ABST
Patent Text Reader

Abstract

The present application discloses an object tracking method, device, equipment and medium based on pseudo multi-modal retrieval, which relates to the technical field of image processing. The method includes obtaining a search image and a template image; inputting the search image and the template image into an object tracking model, and outputting an object tracking result. Among them, the object tracking model is used to: respectively extract features from the search image and the template image to obtain search features and template features; map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; filter the search features through a progressive dynamic visual feature selection method; perform pseudo multi-modal retrieval and object prediction on the quantized template features and the filtered search features to obtain the object tracking result. The present application improves the accuracy and efficiency of object tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular, to an object tracking method, device, equipment and medium based on pseudo multi-modal retrieval. Background Art

[0002] Object tracking is a core research direction in the field of computer vision, and has received extensive attention due to its important value in multiple key application scenarios such as security monitoring, autonomous driving, human-computer interaction, and sports event analysis. From the perspective of technical classification, object tracking is mainly divided into two categories: single-object tracking and multi-object tracking. In terms of generalization ability and practical application value, single-object tracking has significant advantages over multi-object tracking. Specifically, single-object tracking shows stronger adaptability to object types and can effectively track specific objects in various application scenarios; at the same time, the methodological system of single-object tracking is relatively more concise, which is convenient for implementation and further optimization.

[0003] Although significant progress has been made in academic research and technical applications of object tracking algorithms in recent years, in actual application scenarios, this technology still faces many complex and challenging problems, which seriously restrict its robustness and practicality in dynamic and complex environments. First, when the object is in a high-speed motion state, the position of the object between two consecutive frames of images may shift significantly, resulting in a significant decrease in tracking accuracy. This problem is particularly prominent in high-speed scenarios such as autonomous driving and drone navigation. Second, when there are interfering objects with similar visual features to the object in the scene, it is easy to cause misjudgment or deviation of object positioning. This phenomenon is particularly common in crowded people or complex backgrounds. In addition, when the object is partially or completely occluded, the lack of effective visual information will directly affect the accuracy and continuity of the tracking model. A more complex situation is that when the object briefly leaves the camera's field of view and then enters again, traditional algorithms may not be able to achieve continuous tracking due to incorrect intermediate judgments. These problems not only affect the accuracy of object tracking, but also limit its wide deployment in actual applications. Therefore, further research and technological innovation are crucial for improving the practicality and robustness of object tracking technology. Summary of the Invention

[0004] The purpose of the present application is to provide an object tracking method, device, equipment and medium based on pseudo multi-modal retrieval, so as to improve the accuracy and efficiency of object tracking.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides an object tracking method based on pseudo multi-modal retrieval, and the object tracking method based on pseudo multi-modal retrieval includes:

[0007] Obtain a search image and a template image;

[0008] Input the search image and the template image into a target tracking model, and output a target tracking result, where the target tracking result is to identify a target to be tracked in the search image;

[0009] Among them, the target tracking model is used for:

[0010] Extract features from the search image and the template image respectively to obtain search features and template features;

[0011] Map the template features to the codewords in the codebook, and the mapped codewords form the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder;

[0012] Filter the search features by a progressive dynamic visual feature selection method;

[0013] Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result.

[0014] In a second aspect, the present application provides a target tracking device based on pseudo-multi-modal retrieval. The target tracking device based on pseudo-multi-modal retrieval applies the target tracking method based on pseudo-multi-modal retrieval described in any one of the above. The target tracking device based on pseudo-multi-modal retrieval includes:

[0015] An image acquisition module for obtaining a search image and a template image;

[0016] A target tracking module for inputting the search image and the template image into a target tracking model and outputting a target tracking result, where the target tracking result is to identify a target to be tracked in the search image;

[0017] Among them, the target tracking model is used for:

[0018] Extract features from the search image and the template image respectively to obtain search features and template features;

[0019] Map the template features to the codewords in the codebook, and the mapped codewords form the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder;

[0020] Filter the search features by a progressive dynamic visual feature selection method;

[0021] Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result.

[0022] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the target tracking method based on pseudo multi-modal retrieval described in any one of the above.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the target tracking method based on pseudo multi-modal retrieval described in any one of the above.

[0024] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:

[0025] The present application provides a target tracking method, device, equipment and medium based on pseudo multi-modal retrieval. Feature extraction is respectively performed on a search image and a template image to obtain a search feature and a template feature; a codebook corresponding to the template image is obtained; the codebook is constructed by using a vector quantization variational autoencoder, which maps a continuous high-dimensional feature space to a discrete low-dimensional codebook space, thereby explicitly capturing the high-level semantic information of the target, enhancing the semantic expression ability of the features while removing redundant information, and improving the target tracking accuracy and efficiency; in addition, the template feature is mapped to a codeword in the codebook, and the mapped codewords constitute the quantized template feature; the search feature is filtered by a progressive dynamic visual feature selection method, irrelevant background information and redundant features are removed, the computational complexity is reduced, and the target tracking efficiency is further improved. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0027] Figure 1 It is a schematic flowchart of a target tracking method based on pseudo multi-modal retrieval provided by an embodiment of the present application.

[0028] Figure 2 It is a schematic diagram of the principle of discrete visual semantic codebook representation provided by an embodiment of the present application.

[0029] Figure 3 It is an overall framework diagram of a target tracking model provided by an embodiment of the present application.

[0030] Figure 4 It is a schematic diagram of the structure of the first Transformer layer provided by an embodiment of the present application.

[0031] Figure 5 The structural schematic diagram of a computer device provided by an embodiment of the present application. Specific embodiments

[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0033] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0034] The present application provides a target tracking method based on pseudo-multi-modal retrieval, as Figure 1 shown, the target tracking method based on pseudo-multi-modal retrieval includes step 101-step 102.

[0035] Step 101: Obtain a search image and a template image.

[0036] Step 102: Input the search image and the template image into the target tracking model, and output a target tracking result, where the target tracking result is to identify the target to be tracked in the search image.

[0037] Among them, the target tracking model is used to: respectively extract features from the search image and the template image to obtain search features and template features; map the template features to the codewords in the codebook, and the mapped codewords form the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; filter the search features through a progressive dynamic visual feature selection method; perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result.

[0038] Identifying the target to be tracked in the search image specifically identifies the target to be tracked through a target bounding box.

[0039] In an exemplary embodiment, the target tracking model includes a feature extraction module, and the feature extraction module includes a plurality of first Transformer layers connected in sequence, and each of the first Transformer layers corresponds to a predefined ratio. The feature extraction module of the present application is a feature extraction module that combines joint feature extraction and progressive dynamic visual feature extraction. Among them, the search image and the template image are jointly input into the feature extraction module. Among them, the search features will be progressively filtered (i.e., selected) at a filtering ratio through each layer, that is, after passing through each layer, the sequence length will become shorter to eliminate irrelevant background information and redundant features.

[0040] Joint feature extraction means that the joint embedding of the template features and the search features is input into the feature extraction module, but the computational complexity increases due to the increase in the input sequence length. Progressive means that in each first Transformer layer, some features with a relatively low similarity to the current template are eliminated at a predefined ratio. Specifically, a binary matrix B is obtained for each layer, corresponding to whether each feature is retained or eliminated.

[0041] Filter the search features through the progressive dynamic visual feature selection method, which specifically includes: obtaining the nth layer search features output by the nth first Transformer layer; n is an integer greater than 0; for the nth layer search features, calculate the similarity between the nth layer search features and the template features, and eliminate the nth layer search features with low similarity according to the predefined ratio corresponding to the nth first Transformer layer. For example, sort the similarities from low to high, and eliminate the number of nth layer search features corresponding to the first predefined ratio in the sorting. If the nth first Transformer layer is the last first Transformer layer, then the remaining nth layer search features after elimination are used as the filtered search features, otherwise the remaining nth layer search features after elimination are input into the (n + 1)th first Transformer layer. The structure of the first Transformer layer is as Figure 4 shown Figure 4 where template embedding refers to template feature embedding, search embedding refers to search feature embedding. The template embedding and the search embedding pass through the first layer of normalization (Layer Normalization, LN), multi-head attention (Multi-Head Attention, MSA), the second layer of normalization, and the feedforward neural network (Feedforward Neural Network, FNN) in sequence and enter the progressive dynamic visual feature selection module. Among them, the input of the first layer of normalization is also connected to the output of the multi-head attention, and the input of the second layer of normalization is also connected to the output of the feedforward neural network.

[0042] In an exemplary embodiment, a codebook is constructed using a vector quantization variational autoencoder, specifically including:

[0043] The codebook construction neural network is trained using a codebook construction dataset to obtain the codebook; the codebook construction dataset includes multiple sample images, and the codebook construction neural network includes an encoder and a first decoder; the encoder is used to extract multiple feature vectors from the sample images and map each feature vector to the corresponding codeword in the codebook according to the nearest neighbor principle, and the first decoder is used to perform a decoding operation on the codewords output by the encoder to output a decoded image; the goal of training the codebook construction neural network is to minimize the difference between the sample image and the corresponding decoded image. The codebook construction dataset is a large-scale dataset.

[0044] The target tracking model is obtained by training a target neural network. During the training of the target neural network, the codebook participates in the backpropagation algorithm as an adjustable parameter of the target neural network.

[0045] That is to say, there are mainly two ways to construct the codebook: one is the codebook generated by the pre-training method, that is, generated by training the codebook construction neural network; the other is the codebook that is gradually iteratively optimized during the training of the target tracking model. Specifically, the codebook generated in the pre-training stage mainly relies on the support of a large-scale dataset. In this stage, the encoder is used to extract a series of feature vectors from the original input sample images and map these features to the corresponding codewords in the codebook according to the "nearest neighbor". Subsequently, the second decoder is used to perform a decoding operation on these codewords, and its pre-training goal is to minimize the difference between the decoded image and the original input sample image. This process aims to ensure that the selected codewords can accurately reflect the key features of the original input sample images, thereby indirectly measuring the quality of the codebook. As for the codebook that is continuously iteratively optimized during training, based on the pre-training, it is further fine-tuned for the dataset required for a specific tracking task. During this process, the codebook is regarded as an adjustable parameter of the entire target tracking model and participates in the backpropagation algorithm, and its goal is to maximize the tracking performance index.

[0046] The template image is mapped to a visual codebook constructed based on a vector quantization variational autoencoder, specifically by finding the "nearest neighbor" in the codebook, that is, mapping each template feature to a certain codeword in the codebook; the quantized template features and the filtered visual features (search features) are input into the second decoder, that is, the pseudo-multi-modal retrieval plus prediction head module, and after the attention operation of the Transformer, the final result is output.

[0047] As an efficient visual concept encoding tool, the codebook can represent different visual concepts in a lower dimension and classify similar image patches into representative codewords in the codebook, such as Figure 2 shown. By extracting the basic features of the image, this method not only significantly improves the efficiency of feature extraction, but also enhances the robustness and generalization ability of the features, making the tracking algorithm perform more stably and accurately when dealing with target appearance changes and complex situations (such as motion blur). The core function of the codebook is to establish a mapping mechanism between visual concepts and their corresponding representative codewords in the codebook, such as Figure 2 shown. In the vocabulary, that is, in the word embedding layer, in the input sentence "man surfing on the sea", man is mapped to codeword 1, surf is mapped to codeword 256, and sea is mapped to codeword K. This mapping mechanism can be achieved through pre-training or continuously updated through iterative optimization during the actual training process. This continuous update mechanism ensures that semantically or visually related objects (such as "carnation" and "rose") are closer in the vector space, thus improving the accuracy of semantic similarity. The dynamic adjustment of the codebook is crucial for enhancing the model's ability to recognize and associate similar visual concepts in the embedding space.

[0048] When creating a codebook, if there is a series of feature vectors , where each feature vector belongs to one of the nearest vectors to the codeword in the codebook. In an ideal situation, the update method of the codeword should be the arithmetic mean of all feature vectors in the feature vector set to minimize the average distance between and all vectors in . Its update formula is: .

[0049] However, since the target tracking model is usually trained on small batches of data and cannot obtain all vectors in the feature vector set at the same time, an exponential moving average (EMA) update strategy is adopted to balance the influence of historical state and current state. The EMA update rule dynamically adjusts the update process of the codeword by introducing a decay coefficient to ensure the smoothness and gradualness of the update steps.

[0050] In an exemplary embodiment, the codebook is updated during the training process of the codebook construction neural network using the codebook construction dataset.

[0051] During the process of updating the codebook, for a codeword Update of

[0052] First, update the count of codewords to reflect the number of feature vectors related to in the current batch. The update formula is:

[0053]

[0054] 2) Update the cumulative feature vectors of codewords: Update the cumulative feature vectors of codewords to reflect the sum of feature vectors related to in the current batch. The update formula is:

[0055]

[0056] 3) Update the codewords: Calculate the codewords at the current moment according to the updated count and cumulative feature vectors :

[0057]

[0058] where represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the previous update moment, represents the number of feature vectors related to in the set of feature vectors corresponding to the sample image at the current update moment, is the attenuation coefficient, represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the current update moment, represents the cumulative feature vectors of the codeword at the previous update moment, represents the cumulative feature vectors of the codeword at the current update moment, is a constant, represents the i-th feature vector among the nearest feature vectors to the codeword in the set of feature vectors corresponding to the sample image at the current update moment, is the updated value of the codeword The feature vectors related to the codeword refer to at least one feature vector closest to the codeword in the corresponding set of feature vectors, that is, the feature vectors related to the codeword All the closest eigenvectors, where "closest" means the highest similarity, and similarity can be determined by distance, with the shorter the distance, the higher the similarity.

[0059] Through the above steps, the EMA update strategy can gradually optimize the representation of the codewords while ensuring a smooth update process, thereby enhancing the model's ability to recognize and associate visual concepts. A higher value can further ensure the stability of the update process and avoid fluctuations caused by small batches of data.

[0060] Refinement of the steps for the codebook, i.e., the feature vector quantization part: The codebook is the reference in feature vector quantization, and the most similar codeword is found in this codebook with a limited length.

[0061] After the above EMA update process, the template features after joint feature extraction can be quantized and mapped to the codebook where represents the th codeword, is the size of the visual codebook, and is the dimension of each codeword. represents a - dimensional vector, represents a - dimensional vector, represents the number of codewords in the codebook.

[0062] In an exemplary embodiment, the codebook of the present application, i.e., the visual codebook as a repository of visual markers, can convert continuous feature representations into a set of discrete codewords. However, since the hidden dimension of the Vision Transformer (ViT) backbone network is usually much larger than the dimension of the visual codebook Therefore, it is necessary to first project the template features onto a low - dimensional space aligned with the dimension of the visual codebook through a Multi - Layer Perceptron (MLP)

[0063] Mapping the template features to the codewords in the codebook specifically includes:

[0064] Projecting each initial feature vector in the template features onto the space aligned with the dimension of the codebook through a multi - layer perceptron to obtain each feature vector (the projected feature vector), and this projection process can be expressed as:

[0065]

[0066] where​​ is the set of projected feature vectors.

[0067] The core of the quantization operation in this application is to assign each projected feature vector to the closest codeword in the codebook . To achieve this goal, the feature vectors and codebook entries are first normalized.

[0068] Assign a closest codeword from the codebook to each feature vector, specifically including: assign a closest codeword to each feature vector according to the following formula:

[0069]

[0070]

[0071] where, is the codeword of the i-th feature vector, is the -th codeword in the codebook, is the index of the codeword closest to the i-th feature vector, represents the i-th feature vector among the feature vectors closest to the codeword in the template features, represents the k-th codeword in the codebook, represents the L2 norm of the vector.

[0072] Through the above process, the continuous feature space is converted into a discrete set of codewords , is the number of codewords, is the n-th codeword in the codebook. The quantized discrete set of codewords and the features of the search image are input into the second decoder for subsequent pseudo multi-modal retrieval process.

[0073] In an exemplary embodiment, the progressive dynamic visual feature selection method of this application is specifically a progressive dynamic visual feature selection method based on joint feature extraction. This method specifically includes: after the template image and the search image are cut, linearly projected, and superimposed with position embedding (tensor addition of template features and search features), the initial image features and are obtained respectively.

[0074] where, is the template feature ( Figure 3 in ), is the search feature ( Figure 3 in ), represents A vector of dimension

[0075] The progressive dynamic vision feature selection method based on joint feature extraction uses prior knowledge to evaluate the similarity between each search feature and the template feature, and only retains the most informative markers for subsequent processing, thereby effectively reducing the computational overhead and improving the feature quality.

[0076] Specifically, a binary matrix is introduced to indicate whether each search feature is retained or deleted:

[0077]

[0078] where is a constant, is the retention flag of the i-th search feature, 1 means retained, 0 means deleted; is the filtering ratio (predefined ratio), represents the retention score of the i-th search feature, that is, the search feature with the highest similarity to the template feature.

[0079] The progressive dynamic vision feature selection method based on joint feature extraction is essentially a multi-layer Transformer, that is, multiple first Transformer layers. Each first Transformer layer will obtain a binary matrix B, corresponding to which features with lower similarity are "progressively" eliminated.

[0080] In each layer (the first Transformer layer), first use the binary matrix of the previous layer to process the feature matrix . This ensures that the embeddings discarded in the previous layer are replaced by zero vectors, thus avoiding interference with subsequent feature similarity calculations:

[0081]

[0082] where is the binary matrix of the -th layer, represents element-wise multiplication, is the processed feature matrix .

[0083] In the first layer ( ), all search features are retained, so the binary matrix is initialized as:

[0084]

[0085] where is a length of All - one vector of (the number of search features).

[0086] In each subsequent layer, the binary matrix is updated according to the following steps ( ), L where is the number of the first Transformer layer.

[0087] In an exemplary embodiment, the progressive dynamic visual feature selection method more specifically includes the following steps.

[0088] 1) Calculate the retention score .

[0089] For each search feature , calculate the sum of the similarity scores between it and all template features , which is called the retention score. Specifically, the similarity score between the th search feature and the th template feature is calculated by cosine similarity:

[0090]

[0091] The retention score of the th search feature in the th first Transformer layer is calculated as follows:

[0092]

[0093] where

[0094] is the number of template features. 2) Apply the filtering ratio

[0095] . Determine the top embeddings to be retained according to the filtering ratio (different layers may use different filtering ratios). The binary matrix

[0096] is updated as follows:

[0097] where

[0098] is the indicator function (returns 1 when the condition is true, otherwise returns 0), returns the th largest value of the retention score ​​​​The dot product ensures that the embeddings discarded in the previous layer remain discarded in the current layer.

[0099] 3) Propagate the binary matrix to the subsequent layer .

[0100] Pass the updated binary matrix as input to the next layer ( ) for further feature selection and refinement.

[0101] After processing all layers ( layers) of the first Transformer layer, the final binary matrix is obtained (obtaining means obtaining which features of the search image are finally retained through the "Progressive Dynamic Visual Feature Selection Method Based on Joint Feature Extraction", Figure 3 the part represented by the dashed features in Figure 3 is filtered and not selected),

[0102] Among them, L stacking the first Transformer layers is the "Joint Feature Extraction + Progressive Visual Feature Selection" in the figure, and each layer "progressively" filters and eliminates some features through a filtering ratio.

[0103] Among them, Q X , K Z , V Z These parameters are very standard Transformer operations. Q X , K Z , V Z are the query, key, and value matrices after linear transformation respectively. K Z and V Z are obtained through different linear mappings of the quantized template features, and Q X is obtained through linear mapping of the search features after "visual feature selection".

[0104] By introducing the progressive dynamic visual feature selection method, this application not only significantly reduces the interference of irrelevant background information on feature learning, but also greatly reduces the computational complexity, thereby improving the training and inference efficiency of the model while ensuring the tracking accuracy. This method provides an efficient and robust solution for object tracking tasks in complex scenarios.

[0105] In an exemplary embodiment, the target tracking model includes a second decoder, which is used to perform pseudo-multimodal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result.

[0106] In the target tracking task, image templates are usually used to describe the appearance features of the target. In this application, these image templates are regarded as a special kind of "text prompt", and by classifying and mapping similar image patches into representative codewords, low-dimensional semantic coding is achieved. This discretized representation method can not only capture target features with higher accuracy and stronger stability, discard complex details and irrelevant information in the image, but also convey core semantic content in a more abstract way, while avoiding the complexity of natural language. Pseudo-multimodality significantly improves the adaptability of the target tracking algorithm in a dynamic environment, especially in the case where the target moves rapidly or is occluded, showing stronger stability and accuracy.

[0107] After feature extraction and vector quantization in this application, a second decoder is used to implement the pseudo-multimodal retrieval task. The second decoder is composed of multiple stacked second Transformer layers. That is, the second decoder includes a plurality of second Transformer layers connected in sequence, and each of the second Transformer layers includes a multi-head cross-attention module (MCA). Among them, the MCA module is used to promote the interaction between the abstract template features and the search image features to achieve refined modeling of the target features. This iterative retrieval process can be represented by the following formula:

[0108]

[0109]

[0110] Among them, represents the initial search image features, is the number of second Transformer layers, represents the intermediate features output by the

[0111] The motivation for introducing the cross-attention mechanism is to explicitly model the relationship between two sets of different features: the discrete codeword set and the search image features 。Different from the traditional self-attention mechanism (which concatenates two sets of features into a single token set and calculates internal relationships), the cross-attention mechanism directly calculates the attention weights between two sets of tokens, enabling the model to more specifically focus on the regions in the search image that are most relevant to the target. Additionally, since the target features are represented in a discrete and quantized form, lacking high-frequency details, they can better adapt to changes in the target appearance (such as scale, rotation, or illumination changes).

[0112] Specifically, the calculation process of the cross-attention mechanism is as follows:

[0113]

[0114]

[0115] Among them, 、 and represent the query, key, and value matrices respectively, 、 and are the corresponding learnable weight matrices, is the length of the key matrix, and T represents the transpose.

[0116] Finally, the refined features processed by the layer of the second Transformer layer are passed to the prediction head to generate the final prediction results for the target location.

[0117] This design not only enhances the modeling ability for target features but also improves the robustness and generalization performance of the model in complex scenarios. The final target bounding box is calculated as follows:

[0118]

[0119] Among them, in the target classification score map representing the probability of the target's existence at each spatial position, the position with the highest score is regarded as the center position of the target, and its calculation formula is:

[0120]

[0121] Among them, is the probability of the target's existence at position in the score map , the local offset map is used to compensate for the discretization error introduced due to the reduction in the feature map resolution, and a normalized bounding box size map specifies the width and height of the bounding box at each position, represents the height of the search image, Represents the width of the search image.

[0122] This application discloses the following technical effects:

[0123] This application first constructs a visual codebook using a vector quantization variational autoencoder, mapping the continuous high-dimensional feature space to a discrete low-dimensional codebook space, thereby explicitly capturing the high-level semantic information of the target. This process not only enhances the semantic expression ability of the features but also improves the robustness of the model by removing redundant information. Subsequently, this application introduces a progressive dynamic visual feature selection method, which dynamically evaluates the similarity between each search feature and the template feature, gradually eliminating irrelevant background information and redundant features, thereby significantly reducing the computational complexity and enhancing the model's ability to focus on the core features of the target. Finally, this application proposes a target tracking paradigm based on pseudo-multimodal retrieval. By matching the image template mapped to the discrete low-dimensional codebook space with the search image in the semantic space, only the core semantic information required for target tracking is retained, and complex and irrelevant details are discarded, thereby significantly improving the accuracy and efficiency of target tracking. This application not only provides an innovative solution for target tracking technology in complex scenarios but also brings new impetus to the research and application in the field of computer vision.

[0124] Based on the same inventive concept, the embodiments of this application also provide a pseudo-multimodal retrieval-based target tracking device for implementing the above-mentioned pseudo-multimodal retrieval-based target tracking method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the pseudo-multimodal retrieval-based target tracking device can refer to the limitations of the pseudo-multimodal retrieval-based target tracking method in the above text and will not be elaborated here.

[0125] In an exemplary embodiment, a pseudo-multimodal retrieval-based target tracking device is provided. The pseudo-multimodal retrieval-based target tracking device applies the above-mentioned pseudo-multimodal retrieval-based target tracking method. The pseudo-multimodal retrieval-based target tracking device includes:

[0126] An image acquisition module for acquiring a search image and a template image.

[0127] A target tracking module for inputting the search image and the template image into a target tracking model and outputting a target tracking result, where the target tracking result is to identify the target to be tracked in the search image.

[0128] Wherein, the target tracking model is used to:

[0129] Extract features from the search image and the template image respectively to obtain search features and template features.

[0130] Map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder.

[0131] Filter the search features by a progressive dynamic visual feature selection method.

[0132] Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result.

[0133] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store target tracking data based on pseudo-multi-modal retrieval. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a target tracking method based on pseudo-multi-modal retrieval.

[0134] Those skilled in the art can understand that Figure 5 the structure shown in

[0135] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, it implements the steps in the above-mentioned method embodiments.

[0136] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memories can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0138] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, data processing logics of programmable logics, etc., and are not limited thereto.

[0139] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0140] In this text, specific examples are used to illustrate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A target tracking method based on pseudo multi-modal retrieval, characterized in that The target tracking method based on pseudo multimodal retrieval includes: Get the search image and the template image; Inputting the search image and the template image into a target tracking model, and outputting a target tracking result, wherein the target tracking result is a target to be tracked identified in the search image; Wherein, the target tracking model is used for: Extracting features from the search image and the template image respectively to obtain search features and template features; Mapping the template features to codewords in a codebook, each mapped codeword constitutes a quantized template feature; the codebook is constructed using a vector quantization variational autoencoder; filtering the search features by a progressive dynamic visual feature selection method; The quantized template features and the filtered search features are used for pseudo multimodal retrieval and target prediction to obtain the target tracking result; The target tracking model includes a feature extraction module, wherein the feature extraction module includes a plurality of first Transformer layers connected in sequence, each of the first Transformer layers corresponding to a predefined ratio; The search features are filtered by a progressive dynamic visual feature selection method, specifically including: Get the n-th layer search feature output by the n-th first Transformer layer; n is an integer greater than 0; For the n-th layer search feature, calculate the similarity between the n-th layer search feature and the template feature, and remove the n-th layer search feature with low similarity according to the predefined ratio corresponding to the n-th first Transformer layer. If the n-th first Transformer layer is the last first Transformer layer, the remaining n-th layer search feature after removal is used as the filtered search feature, otherwise the remaining n-th layer search feature after removal is input into the n+1-th first Transformer layer; The codebook is constructed using a vector quantized variational autoencoder, which includes: The codebook is constructed by training a codebook-constructed neural network using a codebook-constructed data set to obtain the codebook; the codebook-constructed data set includes a plurality of sample images, and the codebook-constructed neural network includes an encoder and a first decoder; the encoder is used to extract a plurality of feature vectors from the sample image, and map each feature vector to a corresponding codeword in the codebook according to the nearest neighbor principle, and the first decoder is used to decode the codeword output by the encoder and output a decoded image; the goal of training the codebook-constructed neural network is to minimize the difference between the sample image and the corresponding decoded image; The target tracking model is obtained by training the target neural network. During the training of the target neural network, the codebook participates in the back propagation algorithm as an adjustable parameter of the target neural network. Mapping the template features to codewords in a codebook specifically includes: Projecting each initial feature vector in the template feature to the space aligned with the codebook dimension through a multi-layer perceptron to obtain each feature vector; Each feature vector is assigned a codeword that is closest to it from the codebook.

2. The target tracking method based on pseudo-multimodal retrieval according to claim 1, characterized in that, In the process of training the neural network constructed by the codebook using the codebook constructed data set, the codebook is updated; During the process of updating the codebook, for a codeword , the update formula is expressed as: ; ; ; Among them, represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the previous update time, represents the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the current update time, is the attenuation coefficient, represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the current update time, represents the cumulative feature vector of the codeword at the previous update time, represents the cumulative feature vector of the codeword at the current update time, is a constant, represents the i-th feature vector among the nearest feature vectors in the set of feature vectors corresponding to the sample image at the current update time, is the updated value of the codeword , and the feature vectors related to the codeword refer to at least one feature vector closest to the codeword in the corresponding set of feature vectors.

3. The target tracking method based on pseudo-multimodal retrieval according to claim 1, wherein, Assign a codeword closest to each feature vector from the codebook, specifically including: Assign a codeword closest to each feature vector according to the following formula: ; ; Among them, is the codeword of the i-th feature vector, is the -th codeword in the codebook, is the index of the codeword closest to the i-th feature vector, represents the i-th feature vector among the closest feature vectors to the codeword, represents the k-th codeword in the codebook.

4. The target tracking method based on pseudo-multimodal retrieval according to claim 1, characterized in that, The target tracking model includes a second decoder, and the second decoder is used to perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result; The second decoder includes a plurality of second Transformer layers connected in sequence, and each second Transformer layer includes a multi-head cross-attention module.

5. An object tracking device based on pseudo multi-modal retrieval, characterized in that, The target tracking device based on pseudo-multi-modal retrieval applies the target tracking method based on pseudo-multi-modal retrieval according to any one of claims 1-4. The target tracking device based on pseudo-multi-modal retrieval includes: An image acquisition module for acquiring a search image and a template image; A target tracking module for inputting the search image and the template image into the target tracking model and outputting the target tracking result, where the target tracking result is to identify the target to be tracked in the search image; Wherein, the target tracking model is used for: Extract features from the search image and the template image respectively to obtain search features and template features; Map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; Filter the search features by a progressive dynamic visual feature selection method; Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result.

6. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the target tracking method based on pseudo-multi-modal retrieval according to any one of claims 1-4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the target tracking method based on pseudo-multi-modal retrieval according to any one of claims 1-4.

Citation Information

Patent Citations

  • Autoregressive visual target tracking algorithm based on token fusion

    CN119477974A

  • Visual target tracking method and device based on efficient sequence generation and electronic equipment

    CN119785355A