Target tracking method and device based on pseudo-multi-mode retrieval, equipment and medium
By adopting pseudo-multimodal retrieval and visual codebook technology in the target tracking technology, the high-level semantic information of the target is explicitly captured, and the features are optimized through the progressive dynamic visual feature selection method, the accuracy problem of target tracking in high-speed motion and complex background is solved, achieving a more efficient and robust target tracking effect.
Patent Information
- Application Number
- CN202510629024.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
Existing target tracking technologies perform poorly in high-speed motion and complex backgrounds, especially in high-speed scenarios and dense crowds or complex backgrounds, which can easily lead to reduced tracking accuracy and misjudgment.
The target tracking method based on pseudo-multimodal retrieval is adopted to construct a visual codebook through a vector quantization variational autoencoder, map the high-dimensional feature space to the low-dimensional codebook space, explicitly capture the high-level semantic information of the target, and filter the search features through the progressive dynamic visual feature selection method to eliminate irrelevant background information.
It improves the accuracy and efficiency of target tracking, enhances the semantic expression ability of features, reduces the computational complexity, improves the robustness and generalization ability of the model, and can handle target appearance changes and complex scenarios more stably and accurately.
Smart Images

Figure CN120147366A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and particularly to an object tracking method, device, equipment and medium based on pseudo multi-modal retrieval. Background Art
[0002] Object tracking is a core research direction in the field of computer vision, and has received extensive attention due to its important value in multiple key application scenarios such as security monitoring, autonomous driving, human-computer interaction, and sports event analysis. From the perspective of technical classification, object tracking is mainly divided into two categories: single-object tracking and multi-object tracking. In terms of generalization ability and practical application value, single-object tracking has significant advantages over multi-object tracking. Specifically, single-object tracking shows stronger adaptability to object types and can effectively track specific objects in a variety of application scenarios; at the same time, the methodological system of single-object tracking is relatively more concise, which is convenient for implementation and further optimization.
[0003] Although significant progress has been made in academic research and technical applications of object tracking algorithms in recent years, in actual application scenarios, this technology still faces many complex and challenging problems, which seriously restrict its robustness and practicality in dynamic and complex environments. First, when the object is in a high-speed motion state, the position of the object between two consecutive frames of images may shift significantly, resulting in a significant decrease in tracking accuracy. This problem is particularly prominent in high-speed scenarios such as autonomous driving and unmanned aerial vehicle navigation. Second, when there are interfering objects with similar visual features to the object in the scene, it is easy to cause misjudgment or deviation of object positioning. This phenomenon is particularly common in crowded people or complex backgrounds. In addition, when the object is partially or completely occluded, the lack of effective visual information will directly affect the accuracy and continuity of the tracking model. A more complex situation is that when the object briefly leaves the camera's field of view and then enters again, traditional algorithms may not be able to achieve continuous tracking due to incorrect intermediate judgments. These problems not only affect the accuracy of object tracking, but also limit its wide deployment in actual applications. Therefore, further research and technological innovation are crucial for improving the practicality and robustness of object tracking technology. Summary of the Invention
[0004] The purpose of the present application is to provide an object tracking method, device, equipment and medium based on pseudo multi-modal retrieval to improve the accuracy and efficiency of object tracking.
[0005] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides an object tracking method based on pseudo multi-modal retrieval, and the object tracking method based on pseudo multi-modal retrieval includes: Obtain a search image and a template image; Input the search image and the template image into the target tracking model, and output the target tracking result, where the target tracking result is to identify the target to be tracked in the search image; Among them, the target tracking model is used to: Extract features from the search image and the template image respectively to obtain search features and template features; Map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; Filter the search features through a progressive dynamic visual feature selection method; Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result.
[0006] In a second aspect, the present application provides a target tracking device based on pseudo-multi-modal retrieval. The target tracking device based on pseudo-multi-modal retrieval applies the target tracking method based on pseudo-multi-modal retrieval described in any one of the above. The target tracking device based on pseudo-multi-modal retrieval includes: An image acquisition module for acquiring a search image and a template image; A target tracking module for inputting the search image and the template image into the target tracking model and outputting the target tracking result, where the target tracking result is to identify the target to be tracked in the search image; Among them, the target tracking model is used to: Extract features from the search image and the template image respectively to obtain search features and template features; Map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; Filter the search features through a progressive dynamic visual feature selection method; Perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain the target tracking result.
[0007] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the steps of the target tracking method based on pseudo-multi-modal retrieval described in any one of the above.
[0008] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the target tracking method based on pseudo-multi-modal retrieval described in any one of the above are implemented.
[0009] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application: This application provides a target tracking method, device, equipment and medium based on pseudo multi-modal retrieval, which respectively extracts features from a search image and a template image to obtain search features and template features; obtains a codebook corresponding to the template image; the codebook is constructed by using a vector quantization variational autoencoder, which maps the continuous high-dimensional feature space to a discrete low-dimensional codebook space, thereby explicitly capturing the high-level semantic information of the target, enhancing the semantic expression ability of the features while removing redundant information, and improving the accuracy and efficiency of target tracking; in addition, mapping the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; filtering the search features by a progressive dynamic visual feature selection method, removing irrelevant background information and redundant features, reducing the computational complexity, and further improving the target tracking efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 It is a schematic flowchart of a target tracking method based on pseudo multi-modal retrieval provided by an embodiment of this application.
[0012] Figure 2 It is a schematic diagram of the principle of discrete visual semantic codebook representation provided by an embodiment of this application.
[0013] Figure 3 It is an overall framework diagram of a target tracking model provided by an embodiment of this application.
[0014] Figure 4 It is a schematic diagram of the structure of the first Transformer layer provided by an embodiment of this application.
[0015] Figure 5 It is a schematic diagram of the structure of a computer device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0017] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] The present application provides a target tracking method based on pseudo-multi-modal retrieval, as Figure 1 shown, the target tracking method based on pseudo-multi-modal retrieval includes step 101-step 102.
[0019] Step 101: Obtain a search image and a template image.
[0020] Step 102: Input the search image and the template image into the target tracking model, and output a target tracking result, where the target tracking result is to identify the target to be tracked in the search image.
[0021] Among them, the target tracking model is used to: respectively extract features from the search image and the template image to obtain search features and template features; map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder; filter the search features through a progressive dynamic visual feature selection method; perform pseudo-multi-modal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result.
[0022] Identifying the target to be tracked in the search image specifically identifies the target to be tracked through a target bounding box.
[0023] In an exemplary embodiment, the target tracking model includes a feature extraction module, and the feature extraction module includes a plurality of first Transformer layers connected in sequence, and each of the first Transformer layers corresponds to a predefined ratio. The feature extraction module of the present application is a feature extraction module that combines joint feature extraction and progressive dynamic visual feature extraction. Among them, the search image and the template image are jointly input into the feature extraction module. Among them, the search features will be progressively filtered (i.e., selected) at a filtering ratio through each layer, that is, the sequence length will become shorter after passing through each layer to eliminate irrelevant background information and redundant features.
[0024] Joint feature extraction means that the joint embedding of the template features and the search features is input into the feature extraction module, but the computational complexity increases due to the increase in the input sequence length. Progressive means that in each first Transformer layer, some features with a low similarity to the current template are eliminated at a predefined ratio. Specifically, a binary matrix B is obtained in each layer, corresponding to whether each feature is retained or eliminated.
[0025] Filter the search features through a progressive dynamic visual feature selection method, specifically including: obtaining the nth layer of search features output by the nth first Transformer layer; n is an integer greater than 0; for the nth layer of search features, calculate the similarity between the nth layer of search features and the template features, and eliminate the nth layer of search features with low similarity according to the predefined ratio corresponding to the nth first Transformer layer. For example, sort the similarities from low to high, and eliminate the number of the nth layer of search features corresponding to the first predefined ratio in the sorting. If the nth first Transformer layer is the last first Transformer layer, then the remaining nth layer of search features after elimination is used as the filtered search features, otherwise the remaining nth layer of search features after elimination is input into the (n + 1)th first Transformer layer. The structure of the first Transformer layer is as Figure 4 shown, Figure 4 where the template embedding refers to the template feature embedding, the search embedding refers to the search feature embedding, and the template embedding and the search embedding sequentially pass through the first layer of layer normalization (LN), multi-head attention (MSA), the second layer of layer normalization, and a feedforward neural network (FNN) to enter the progressive dynamic visual feature selection module. Among them, the input of the first layer of layer normalization is also connected to the output of the multi-head attention, and the input of the second layer of layer normalization is also connected to the output of the feedforward neural network.
[0026] In an exemplary embodiment, a vector quantization variational autoencoder is used to construct a codebook, specifically including: Using a codebook construction dataset to train a codebook construction neural network to obtain the codebook; the codebook construction dataset includes multiple sample images, and the codebook construction neural network includes an encoder and a first decoder; the encoder is used to extract multiple feature vectors from the sample images, map each feature vector to the corresponding codeword in the codebook according to the nearest neighbor principle, and the first decoder is used to perform a decoding operation on the codewords output by the encoder to output a decoded image; the goal of training the codebook construction neural network is to minimize the difference between the sample image and the corresponding decoded image. The codebook construction dataset is a large-scale dataset.
[0027] The target tracking model is obtained by training a target neural network. During the training of the target neural network, the codebook participates in the backpropagation algorithm as an adjustable parameter of the target neural network.
[0028] That is to say, there are mainly two ways to construct the codebook: one is the codebook generated by the pre-training method, that is, training the neural network for codebook construction to generate it; the other is the codebook that is gradually iteratively optimized during the training process of the target tracking model. Specifically, the codebook generated in the pre-training stage mainly relies on the support of large-scale datasets. In this stage, the encoder is used to extract a series of feature vectors from the original input sample images and map these features to the corresponding codewords in the codebook according to the "nearest neighbor". Subsequently, the second decoder is used to decode these codewords, and its pre-training goal is to minimize the difference between the decoded image and the original input sample image. This process aims to ensure that the selected codewords can accurately reflect the key features of the original input sample images, thereby indirectly measuring the quality of the codebook. As for the codebook that is continuously iteratively optimized during training, based on the pre-training, it is further fine-tuned for the dataset required for specific tracking tasks. During this process, the codebook is regarded as an adjustable parameter of the entire target tracking model and participates in the backpropagation algorithm, and its goal is to maximize the tracking performance index.
[0029] The template image is mapped to the visual codebook constructed based on the vector quantization variational autoencoder. Specifically, the "nearest neighbor" is found from the codebook, that is, each template feature is mapped to a certain codeword in the codebook; the quantized template features and the filtered visual features (search features) are input into the second decoder, that is, the pseudo-multi-modal retrieval plus prediction head module, and after the attention operation of the Transformer, the final output result is obtained.
[0030] As an efficient visual concept encoding tool, the codebook can represent different visual concepts in a lower dimension and classify similar image patches into representative codewords in the codebook, such as Figure 2 shown. By extracting the basic features of the image, this method not only significantly improves the efficiency of feature extraction, but also enhances the robustness and generalization ability of the features, making the tracking algorithm perform more stably and accurately when dealing with target appearance changes and complex situations (such as motion blur). The core function of the codebook is to establish a mapping mechanism between visual concepts and their corresponding representative codewords in the codebook, such as Figure 2 shown. In the vocabulary, that is, in the word embedding layer, in the input sentence "man surfing on the sea", man is mapped to codeword 1, surf is mapped to codeword 256, and sea is mapped to codeword K. This mapping mechanism can be achieved through pre-training or continuously updated through iterative optimization during the actual training process. This continuous update mechanism ensures that semantically or visually related objects (such as "carnation" and "rose") are closer in the vector space, thereby improving the accuracy of semantic similarity. The dynamic adjustment of the codebook is crucial for enhancing the model's ability to recognize and associate similar visual concepts in the embedding space.
[0031] When creating a codebook, if there is a series of feature vectors , where each feature vector belongs to one of the vectors that are closest to the codewords in the codebook . In an ideal situation, the codeword should be updated as the arithmetic mean of all the feature vectors in the feature vector set to minimize the average distance between and all the vectors in . The update formula is:
[0032] However, since the target tracking model is usually trained on small batches of data and cannot obtain all the vectors in the feature vector set at the same time, an Exponential Moving Average (EMA) update strategy is adopted to balance the influence of the historical state and the current state. The EMA update rule dynamically adjusts the update process of the codeword by introducing a decay coefficient to ensure the smoothness and gradualness of the update steps.
[0033] In an exemplary embodiment, the codebook is updated during the training process of the codebook construction neural network using the codebook construction dataset.
[0034] During the update process of the codebook, for the update of a codeword : 1) Update the count of the codeword.
[0035] First, update the count of the codeword to reflect the number of feature vectors related to in the current batch. The update formula is:
[0036] 2) Update the cumulative feature vector of the codeword: Update the cumulative feature vector of the codeword to reflect the sum of the feature vectors related to in the current batch. The update formula is:
[0037] 3) Update the codeword: Calculate the codeword at the current moment according to the updated count and cumulative feature vector:
[0038] where, Represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the previous update time, represents the number of feature vectors related to in the set of feature vectors corresponding to the sample image at the current update time, is the decay coefficient, represents the weighted value of the number of feature vectors related to the codeword in the set of feature vectors corresponding to the sample image at the current update time, represents the cumulative feature vectors of the codeword at the previous update time, represents the cumulative feature vectors of the codeword at the current update time, is a constant, represents the i-th feature vector among the nearest feature vectors in the set of feature vectors corresponding to the sample image at the current update time, is the updated value of the codeword , and the feature vectors related to the codeword refer to at least one feature vector in the corresponding set of feature vectors that is closest to the codeword , that is, the feature vectors related to the codeword refer to all feature vectors in the corresponding set of feature vectors that are closest to the codeword . The closest means the highest similarity, and the similarity can be determined according to the distance. The shorter the distance, the higher the similarity.
[0039] Through the above steps, the EMA update strategy can gradually optimize the representation of the codewords while ensuring the smoothness of the update process, thereby enhancing the model's ability to recognize and associate visual concepts. A higher value can further ensure the stability of the update process and avoid fluctuations caused by small batches of data.
[0040] Refinement of the steps for the codebook, that is, the feature vector quantization part: The codebook is a reference in feature vector quantization, and the most similar codeword is found in this codebook with a limited length.
[0041] After the above EMA update process, the template features after joint feature extraction can be quantized and mapped into the codebook , where represents the -th codeword, is the size of the visual codebook, is the dimension of each codeword. represents a -dimensional vector, denote a vector of dimension indicating the number of codewords in the codebook.
[0042] In an exemplary embodiment, the codebook of the present application, i.e., the visual codebook as a repository of visual markers, can convert continuous feature representations into a discrete set of codewords. However, due to the hidden dimension of the Vision Transformer (ViT) backbone is usually much larger than the dimension of the visual codebook Therefore, it is necessary to first project the template features through a Multi-Layer Perceptron (MLP) into a low-dimensional space aligned with the visual codebook dimension
[0043] Mapping the template features to the codewords in the codebook specifically includes: Projecting each initial feature vector in the template features into the space aligned with the codebook dimension through a multi-layer perceptron to obtain each feature vector (the projected feature vector), and this projection process can be expressed as:
[0044] where is the set of projected feature vectors.
[0045] The core of the quantization operation of the present application is to assign each projected feature vector to the closest codeword in the codebook . To achieve this goal, the feature vectors and codebook entries are first normalized.
[0046] Assigning a closest codeword in the codebook to each feature vector specifically includes: Assigning a closest codeword to each feature vector according to the following formula:
[0047]
[0048] where is the codeword of the i-th feature vector, is the -th codeword in the codebook, is the index of the closest codeword to the i-th feature vector, represents the i-th feature vector among the closest feature vectors to the codeword represents the k-th codeword in the codebook, Represents the L2 norm of the vector.
[0049] Through the above process, the continuous feature space is converted into a set of discrete codewords , where is the number of codewords,
[0050] In an exemplary embodiment, the progressive dynamic visual feature selection method of the present application is specifically a progressive dynamic visual feature selection method based on joint feature extraction. The method specifically includes: after the template image and the search image are cut, linearly projected, and superimposed with position embeddings (tensor addition of template features and search features), initial image features are obtained respectively and .
[0051] Among them, is the template feature ( Figure 3 in ), is the search feature ( Figure 3 in ), represents -dimensional vector.
[0052] The progressive dynamic visual feature selection method based on joint feature extraction uses prior knowledge to evaluate the similarity between each search feature and the template feature, and only retains the most informative markers for subsequent processing, thereby effectively reducing the computational overhead and improving the feature quality.
[0053] Specifically, a binary matrix is introduced to represent whether each search feature is retained or deleted:
[0054] Among them, is a constant, is the retention flag of the i-th search feature, 1 means retained, 0 means deleted; is the filtering ratio (predefined ratio), represents the retention score of the i-th search feature, that is, the search feature with the highest similarity to the template feature.
[0055] The progressive dynamic visual feature selection method based on joint feature extraction is essentially a multi-layer Transformer, namely multiple first Transformer layers. Each first Transformer layer will obtain a binary matrix B, corresponding to which features with lower similarity are "progressively" removed.
[0056] In each layer (the first Transformer layer), first use the binary matrix of the previous layer to process the feature matrix . This ensures that the embeddings discarded in the previous layer are occupied by zero vectors, thus avoiding interference with subsequent feature similarity calculations:
[0057] where, is the binary matrix of the th layer, represents element-wise multiplication, is the processed feature matrix .
[0058] In the first layer ( ), all search features are retained, so the binary matrix is initialized as:
[0059] where, is a vector of all 1s with a length of (the number of search features).
[0060] In each subsequent layer, the binary matrix is updated according to the following steps ( ), L is the number of the first Transformer layer.
[0061] In an exemplary embodiment, the progressive dynamic visual feature selection method more specifically includes the following steps.
[0062] 1) Calculate the retention score .
[0063] For each search feature , calculate the sum of the similarity scores between it and all template features , which is called the retention score. Specifically, the similarity score between the th search feature and the th template feature is calculated by cosine similarity:
[0064] The retained score of the th search feature in the th first Transformer layer is calculated as follows:
[0065] where is the number of template features.
[0066] 2) Apply the filtering ratio .
[0067] Determine the top embeddings to be retained according to the filtering ratio (different layers may use different filtering ratios). The binary matrix is updated as follows:
[0068] where is the indicator function (returns 1 when the condition is true and 0 otherwise), returns the th largest value of the retained score , and the dot product with ensures that the embeddings discarded in the previous layer are still discarded in the current layer.
[0069] 3) Propagate the binary matrix to the subsequent layer.
[0070] Pass the updated binary matrix as input to the next layer ( ) for further feature selection and refinement.
[0071] After processing all the layers ( layers) of the first Transformer layer, the final binary matrix is obtained (obtaining means obtaining which features of the search image are finally retained by the "Progressive Dynamic Visual Feature Selection Method Based on Joint Feature Extraction", Figure 3 the part represented by the dashed features in Figure 3 is filtered and not selected),
[0072] where LStacking several first Transformer layers forms the "Joint Feature Extraction + Progressive Visual Feature Selection" in the figure. Each layer "progressively" filters and eliminates a part of the features through a filtering ratio.
[0073] Among them, Q X , K Z , V Z These parameters are very standard Transformer operations. Q X , K Z , V Z are the query, key, and value matrices after linear transformation respectively. K Z and V Z are obtained by passing the quantized template features through different linear mappings, and Q X is obtained by passing the search features after "visual feature selection" through a linear mapping.
[0074] By introducing the progressive dynamic visual feature selection method, this application not only significantly reduces the interference of irrelevant background information on feature learning, but also greatly reduces the computational complexity, thus improving the training and inference efficiency of the model while ensuring the tracking accuracy. This method provides an efficient and robust solution for object tracking tasks in complex scenarios.
[0075] In an exemplary embodiment, the object tracking model includes a second decoder, which is used to perform pseudo-multi-modal retrieval and object prediction on the quantized template features and the filtered search features to obtain the object tracking result.
[0076] In object tracking tasks, image templates are usually used to describe the appearance features of the object. This application regards these image templates as a special kind of "text prompt", and realizes low-dimensional semantic coding by classifying and mapping similar image patches into representative codewords. This discrete representation method can not only capture object features with higher accuracy and stronger stability, discard the complex details and irrelevant information in the image, but also convey the core semantic content in a more abstract way, while avoiding the complexity of natural language. Pseudo-multi-modal significantly improves the adaptability of the object tracking algorithm in dynamic environments, especially in the case where the object moves quickly or is occluded, showing stronger stability and accuracy.
[0077] After feature extraction and vector quantization, this application uses a second decoder to implement the pseudo-multi-modal retrieval task. The second decoder is composed of multiple stacked second Transformer layers. That is, the second decoder includes a plurality of second Transformer layers connected in sequence, and each second Transformer layer includes a multi-head cross-attention module (MCA). Among them, the MCA module is used to promote the abstract template features Interaction with search image features to achieve refined modeling of target features. This iterative retrieval process can be represented by the following formula:
[0078]
[0079] where represents the initial search image features, is the number of the second Transformer layer, represents the th intermediate feature output by the second Transformer layer.
[0080] The motivation for introducing the cross-attention mechanism is to explicitly model the relationship between two different sets of features: the discrete codeword set and the search image features . Different from the traditional self-attention mechanism (which concatenates two sets of features into a single token set and calculates the internal relationship), the cross-attention mechanism directly calculates the attention weights between two sets of tokens, enabling the model to more specifically focus on the regions in the search image that are most relevant to the target. In addition, since the target features are represented in a discrete quantization form and lack high-frequency details, they can better adapt to changes in the target appearance (such as scale, rotation, or illumination changes).
[0081] Specifically, the calculation process of the cross-attention mechanism is as follows:
[0082]
[0083] where , and represent the query, key, and value matrices respectively, , and are the corresponding learnable weight matrices, is the length of the key matrix, and T represents the transpose.
[0084] Finally, the refined feature processed by layers of the second Transformer layer is passed to the prediction head to generate the final prediction result for the target location.
[0085] This design not only enhances the modeling ability of target features but also improves the robustness and generalization performance of the model in complex scenarios. The final target bounding box is calculated as follows:
[0086] Among them, in the target classification score map indicating the probability of the target's existence at each spatial position the position with the highest score is regarded as the central position of the target, and its calculation formula is:
[0087] wherein, is the score map at the position the probability of the target's existence, and the local offset map is used to compensate for the discretization error introduced due to the reduction of the feature map resolution. A normalized bounding box size map specifies the width and height of the bounding box at each position, represents the height of the search image, represents the width of the search image.
[0088] The present application discloses the following technical effects: The present application first constructs a visual codebook using a vector quantization variational autoencoder, maps the continuous high-dimensional feature space to a discrete low-dimensional codebook space, thereby explicitly capturing the high-level semantic information of the target. This process not only enhances the semantic expression ability of the features but also improves the robustness of the model by removing redundant information. Subsequently, the present application introduces a progressive dynamic visual feature selection method, which gradually eliminates irrelevant background information and redundant features by dynamically evaluating the similarity between each search feature and the template feature, thereby significantly reducing the computational complexity and enhancing the model's ability to focus on the core features of the target. Finally, the present application proposes a target tracking paradigm based on pseudo-multimodal retrieval. By matching the image template mapped to the discrete low-dimensional codebook space with the search image in the semantic space, only the core semantic information required for target tracking is retained, and complex and irrelevant details are discarded, thereby significantly improving the accuracy and efficiency of target tracking. The present application not only provides an innovative solution for target tracking technology in complex scenarios but also brings new impetus to the research and application in the field of computer vision.
[0089] Based on the same inventive concept, the embodiments of the present application also provide a pseudo-multimodal retrieval-based target tracking device for implementing the above-mentioned pseudo-multimodal retrieval-based target tracking method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more of the following embodiments of the pseudo-multimodal retrieval-based target tracking device can refer to the limitations on the pseudo-multimodal retrieval-based target tracking method in the above text, and will not be elaborated here.
[0090] In an exemplary embodiment, a target tracking device based on pseudo-multimodal retrieval is provided. The target tracking device based on pseudo-multimodal retrieval applies the target tracking method based on pseudo-multimodal retrieval. The target tracking device based on pseudo-multimodal retrieval includes: An image acquisition module, configured to acquire a search image and a template image.
[0091] A target tracking module, configured to input the search image and the template image into a target tracking model, and output a target tracking result, where the target tracking result is to identify a target to be tracked in the search image.
[0092] Wherein, the target tracking model is used for: Extract features from the search image and the template image respectively to obtain search features and template features.
[0093] Map the template features to the codewords in the codebook, and the mapped codewords constitute the quantized template features; the codebook is constructed by using a vector quantization variational autoencoder.
[0094] Filter the search features by a progressive dynamic visual feature selection method.
[0095] Perform pseudo-multimodal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result.
[0096] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as shown in Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store target tracking data based on pseudo-multimodal retrieval. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a target tracking method based on pseudo-multimodal retrieval.
[0097] Those skilled in the art can understand that Figure 5The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0098] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0100] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0101] In each of the embodiments provided in the present application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on a blockchain, etc. In each of the embodiments provided in the present application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a data processing logic unit of a programmable logic device, etc., not limited thereto.
[0102] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0103] Specific examples are used herein to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A target tracking method based on pseudo multimodal retrieval, characterized in that: The target tracking method based on pseudo multimodal retrieval includes: Get the search image and the template image; Inputting the search image and the template image into a target tracking model, and outputting a target tracking result, wherein the target tracking result is a target to be tracked identified in the search image; Wherein, the target tracking model is used for: Extracting features from the search image and the template image respectively to obtain search features and template features; Mapping the template features to codewords in a codebook, each mapped codeword constitutes a quantized template feature; the codebook is constructed using a vector quantization variational autoencoder; filtering the search features by a progressive dynamic visual feature selection method; The quantized template features and filtered search features are used for pseudo-multimodal retrieval and target prediction to obtain the target tracking result.
2. The target tracking method based on pseudo multimodal retrieval according to claim 1, characterized in that: The target tracking model includes a feature extraction module, wherein the feature extraction module includes a plurality of first Transformer layers connected in sequence, each of the first Transformer layers corresponding to a predefined ratio; The search features are filtered by a progressive dynamic visual feature selection method, specifically including: Get the n-th layer search feature output by the n-th first Transformer layer; n is an integer greater than 0; For the n-th layer search features, the similarity between the n-th layer search features and the template features is calculated, and the n-th layer search features with low similarity are eliminated according to the predefined ratio corresponding to the n-th first Transformer layer. If the n-th first Transformer layer is the last first Transformer layer, the remaining n-th layer search features after the elimination are used as the filtered search features, otherwise the remaining n-th layer search features after the elimination are input into the n+1-th first Transformer layer.
3. The target tracking method based on pseudo multimodal retrieval according to claim 1, characterized in that: The codebook is constructed using a vector quantized variational autoencoder, which includes: The codebook is constructed by training a codebook-constructed neural network using a codebook-constructed data set to obtain the codebook; the codebook-constructed data set includes a plurality of sample images, and the codebook-constructed neural network includes an encoder and a first decoder; the encoder is used to extract a plurality of feature vectors from the sample image, and map each feature vector to a corresponding codeword in the codebook according to the nearest neighbor principle, and the first decoder is used to decode the codeword output by the encoder and output a decoded image; the goal of training the codebook-constructed neural network is to minimize the difference between the sample image and the corresponding decoded image; The target tracking model is obtained by training the target neural network. During the training of the target neural network, the codebook participates in the back propagation algorithm as an adjustable parameter of the target neural network.
4. The target tracking method based on pseudo multimodal retrieval according to claim 3, characterized in that: In the process of training the neural network constructed by the codebook using the codebook constructed data set, the codebook is updated; In the process of updating the codebook, for a codeword , the update formula is expressed as: ; ; ; in, Indicates the feature vector set corresponding to the sample image at the last update time and the codeword The weighted value of the number of relevant eigenvectors, Indicates the feature vector set corresponding to the sample image at the current update time and the codeword The number of associated eigenvectors, is the attenuation coefficient, Indicates the feature vector set corresponding to the sample image at the current update time and the codeword The weighted value of the number of relevant eigenvectors, Indicates the last update time code The cumulative eigenvector of Indicates the current update time code The cumulative eigenvector of is a constant, The distance codeword in the feature vector set corresponding to the sample image at the current update time Recent The i-th eigenvector among the eigenvectors, Codeword The updated value, and the codeword The relevant feature vector refers to the feature vector set corresponding to the codeword The closest at least one eigenvector.
5. The target tracking method based on pseudo multimodal retrieval according to claim 1, characterized in that: Mapping the template features to codewords in a codebook specifically includes: Projecting each initial feature vector in the template feature to the space aligned with the codebook dimension through a multi-layer perceptron to obtain each feature vector; Each feature vector is assigned a codeword that is closest to it from the codebook.
6. The target tracking method based on pseudo multimodal retrieval according to claim 5, characterized in that: Allocating a codeword closest to each feature vector from the codebook specifically includes: Assign a codeword closest to each eigenvector according to the following formula: ; ; in, is the codeword of the ith eigenvector, is the first Code words, is the index of the codeword closest to the i-th eigenvector, Represents the distance codeword in the template feature Recent The i-th eigenvector among the eigenvectors, represents the kth codeword in the codebook.
7. The target tracking method based on pseudo multimodal retrieval according to claim 1, characterized in that: The target tracking model includes a second decoder, which is used to perform pseudo multimodal retrieval and target prediction on the quantized template features and the filtered search features to obtain a target tracking result; The second decoder includes a plurality of second Transformer layers connected in sequence, and each of the second Transformer layers includes a multi-head cross attention module.
8. A target tracking device based on pseudo multimodal retrieval, characterized in that: The target tracking device based on pseudo multimodal retrieval applies the target tracking method based on pseudo multimodal retrieval according to any one of claims 1 to 7, and the target tracking device based on pseudo multimodal retrieval comprises: An image acquisition module, used to acquire a search image and a template image; A target tracking module, used to input a search image and a template image into a target tracking model, and output a target tracking result, wherein the target tracking result is a target to be tracked identified in the search image; Wherein, the target tracking model is used for: Extracting features from the search image and the template image respectively to obtain search features and template features; Mapping the template features to codewords in a codebook, each mapped codeword constitutes a quantized template feature; the codebook is constructed using a vector quantization variational autoencoder; filtering the search features by a progressive dynamic visual feature selection method; The quantized template features and filtered search features are used for pseudo-multimodal retrieval and target prediction to obtain the target tracking result.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the target tracking method based on pseudo multimodal retrieval described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target tracking method based on pseudo multimodal retrieval described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image generation method and device, equipment, storage medium and computer program product
CN114638905A
Autoregressive visual target tracking algorithm based on token fusion
CN119477974A
Visual target tracking method and device based on efficient sequence generation and electronic equipment
CN119785355A