Visual target tracking method and device dynamically adapting to target state change

By introducing the attention mechanism of search area guidance in the target tracking method and dynamically updating the characteristics of templates and search areas, the problem of redundancy in the template information and difficulty in adapting to the change of target state in the prior art is solved, and the robustness and accuracy of the tracker are improved.

CN120125616APending Publication Date: 2025-06-10INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325764.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing target tracking method based on attention model is difficult to effectively adapt to the change of target state during the interaction between the template and the search area, resulting in redundant template information and affecting the robustness and accuracy of the tracker.

Method used

The attention mechanism guided by search area is adopted to extract the network through dynamic feature, update the features of the template and search area frame by frame, reduce the impact of irrelevant template information on target representation, and improve the dynamic adaptability of the tracker.

Benefits of technology

It improves the robustness and accuracy of the target tracker under dynamic changes in the target, and reduces the negative impact of template information redundancy on the tracker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125616A_ABST
    Figure CN120125616A_ABST
Patent Text Reader

Abstract

The invention discloses a visual target tracking method and device dynamically adapting to target state changes, and belongs to the field of computer vision. The method comprises the following steps: acquiring a template and a search area image; the dynamic feature extraction network receives the image as input, and maps the image into features embedded with # imgabs0 #, # imgabs1 #, # imgabs2 # and # imgabs3 # as input of an attention mechanism guided by a search area; performing dynamic feature extraction by using an attention mechanism guided by a search area, and outputting template information # imgabs4 # most related to a target state in template features and most important target information # imgabs5 # in search area features; the operation is circularly executed for 11 times, and # imgabs8 # and # imgabs9 # are replaced with # imgabs6 # and # imgabs7 # before execution of each time to serve as input of an attention mechanism guided by a search area; obtaining the final characteristics # imgabs10 # and # imgabs11 # which are most related to the current target state; and finally, predicting a target position according to # imgabs12 #. According to the method, the influence of irrelevant template information on target representation is minimized, and wrong guidance of the irrelevant template information on the tracker is avoided, so that the robustness of the tracker under the dynamic change of the target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a visual target tracking method and device that can dynamically adapt to changes in target state. Background Art

[0002] Target tracking is a fundamental research task in the field of computer vision, whose purpose is to continuously identify and locate the target in a video sequence based on the initial state of the target. Since the target state changes over time, learning a target representation that can dynamically adapt to changes in the target state is crucial for achieving robust visual tracking. To adapt to changes in the target state, existing tracking methods based on the attention model usually learn an accurate representation of the user-specified tracking target through attention interaction between the template and search region information. However, the search region changes dynamically over time and usually contains complex background information. The template image is usually cropped from the first frame of the video sequence, but as the target state changes in subsequent frames, the template may contain some information irrelevant to the target state. In the process of attention interaction between the template and the search region in the prior art, the search region completely refers to all the information of the template to identify the target, and these irrelevant template information will generate responses with similar distractors in the search region, resulting in the enhancement of some incorrect information in the search region, thereby damaging the accuracy of the target representation. Summary of the Invention

[0003] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0004] A visual target tracking method that can dynamically adapt to changes in target state, including:

[0005] Step 1: Obtain a template image and a search region image ;

[0006] Step 2: The dynamic feature extraction network receives the template image and the search region image as inputs, maps the two images into two feature embeddings and , and as inputs to the search region-guided attention mechanism; uses the search region-guided attention mechanism for dynamic feature extraction, so that the template attention branch outputs the template information most relevant to the target state in the template feature , and the search region attention branch outputs the most important target information in the search region feature ;

[0007] Step 3: Repeat Step 2 eleven times. Before each execution, use the output by the template attention branch and the output by the search region attention branch to replace the initial inputs and

[0008] respectively, as the inputs to the search region-guided attention mechanism; and ;

[0009] Step 4: After repeating Step 2 eleven times, the search region-guided attention mechanism obtains and outputs the final features that are most relevant to the current target state

[0010] A visual target tracking device that dynamically adapts to changes in the target state, comprising:

[0011] Image acquisition module: Obtain a template image and a search region image from a given video sequence;

[0012] Dynamic feature extraction network: The dynamic feature extraction network receives the template image and the search region image as inputs, maps the two images into two feature embeddings and using the token transformation module, and as the inputs to the search region-guided attention mechanism; perform dynamic feature extraction using the search region-guided attention mechanism to enable the template attention branch to output the template information in the template features that is most relevant to the target state, and the search region attention branch to output the most important target information in the search region features;

[0013] Loop execution module: Repeat the operations of the dynamic feature extraction network eleven times. Before each execution, use the output by the template attention branch and the output by the search region attention branch to replace the initial inputs and

[0014] respectively, as the inputs to the search region-guided attention mechanism; and ;

[0015] Prediction module: The regression network predicts the target position based on the output in step 4.

[0016] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the visual target tracking method for dynamically adapting to target state changes are implemented.

[0017] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the visual target tracking method for dynamically adapting to target state changes are implemented.

[0018] A computer program product includes a computer program, and when the computer program is executed by a processor, the visual target tracking method for dynamically adapting to target state changes is implemented.

[0019] The present invention has the following beneficial effects:

[0020] The method of the present invention includes an attention mechanism guided by a search region. The attention mechanism aims to first guide the tracker to focus on the most useful template information according to the state of the target in each frame of the search region, and then use the accurate template information to guide the search region to focus on the information most beneficial to identifying and tracking the target. The present invention solves the problems of redundant template information in traditional methods and difficulty in adapting to target state changes, and improves the robustness and accuracy of the tracker in dealing with target dynamic changes. The present invention minimizes the influence of irrelevant template information on target representation, avoids the wrong guidance of the tracker by irrelevant template information, and thus improves the robustness of the tracker under target dynamic changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. is a schematic structural diagram of the visual target tracking method for dynamically adapting to target state changes of the present invention;

[0022] Figure 2 FIG. is a schematic diagram of the attention mechanism guided by a search region. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] The visual object tracking method for dynamically adapting to target state changes of the present invention proposes a novel search region-guided attention mechanism for guiding the tracker to perform dynamic feature extraction according to target state changes during the interaction between the template and the search region. Specifically, the search region-guided attention mechanism is the key module for dynamic feature extraction in the method of the present invention, and it is internally divided into two attention branches, namely the template attention branch and the search region attention branch. The template attention branch receives the query vector of the search region and the query, key, and value vectors of the template as inputs, aiming to dynamically enhance the information in the template that is most relevant to the current target state based on the guidance of the search region query vector, and obtain enhanced template information. The search region attention branch receives the query vector of the search region and the enhanced template information output by the template attention branch as inputs, aiming to dynamically focus on the most important target information in the search region features according to the hint of the enhanced template information. These two attention branches guide each other and are iteratively performed in the dynamic feature extraction network, and finally output a robust target representation for subsequent target position regression prediction.

[0025] This method is divided into three parts: inputting a pair of template and search region images, a dynamic feature extraction network, and using a regression network to predict the target position according to the output of the dynamic feature extraction network.

[0026] For the dynamic feature extraction network: the search region-guided attention mechanism proposed by the present invention is used for dynamic feature extraction, and the initialized weights of the network model adopted include but are not limited to ConvMAE.

[0027] For the pair of input template and search region images, usually the first frame of the sampled test video sequence is used as the template image, and subsequent frames are dynamically sampled as the search region images. Both images contain the same target, the search region contains more background, and the template only contains the target.

[0028] For the regression network, according to the target-related features captured by the search region-guided attention mechanism in the dynamic feature extraction network, a bounding box prediction network (including but not limited to CornerNet) is used to locate the target in the search region, and finally the bounding box of the tracked target is obtained.

[0029] As Figure 1 shown, the specific steps of the visual object tracking method for dynamically adapting to target state changes of the present invention are as follows:

[0030] Step 1: Obtain the template image and the search region image from the given video sequence; the resolution sizes of the template image and the search region image can be set to and Pixel.

[0031] Step 2: The dynamic feature extraction network receives the template image and the search area image as inputs, and maps the two images into two feature embeddings using the token transformation module and , and as the inputs of the search area-guided attention mechanism.

[0032] As Figure 2 shown, the interaction process between the template attention branch and the search area attention branch in the search area-guided attention mechanism in Step 2 is specifically as follows:

[0033] Step 2.1: Unfold and concatenate them along the and spatial dimensions, and map and using a fully connected layer to obtain the query vector representing the template features, the key vector and the value vector , as well as the query vector representing the search area features, the key vector and the value vector .

[0034] Step 2.2: The template attention branch receives the query vector of the search area with the current target state , the key vector and the value vector of the template as inputs.

[0035] The attention interaction process between the query vector , the key vector and the value vector of the template and the query vector of the search area in Step 2.2 is specifically as follows:

[0036] Step 2.2.1: The query vector of the search area first performs matrix calculation with the key vector of the template to obtain the similarity matrix between them. Each row of this matrix represents the similarity between a search area feature and the entire template feature;

[0037] Step 2.2.2: The query vector of the template and the key vector perform matrix calculation to obtain the similarity matrix between all template features , each row of the matrix represents the similarity between a template feature and the entire template feature;

[0038] Step 2.2.3: For the matrix calculate the mean of each row of features, and select the top rows with the largest mean (to filter out the irrelevant attention responses caused by the interference information between the template and the search area), obtaining a new attention matrix in the template attention branch that is most relevant to the target feature , where is set to the number of rows;

[0039] Step 2.2.4: Add and element-wise to introduce the similarity related to the current target state into , obtaining the final similarity matrix in the template attention branch;

[0040] Step 2.2.5: Perform an activation function operation on the matrix to concentrate the weights on the information most relevant to the target, and then perform a matrix multiplication on and the value vector of the template to weight the target information in the value vector of the template according to the similarity weights in the matrix , obtaining the template information in the template feature that is most relevant to the target state .

[0041] Step 2.3: The search area attention branch receives the query vector of the search area with the current target state and the concatenated key vector and value vector of the template and the search area as inputs, where represents the concatenation of two vectors along the spatial dimension, represents any vector.

[0042] The specific process of the attention interaction between the query vector of the search area and the concatenated key vector and value vector of the template and the search area in Step 2.3 is as follows:

[0043] Step 2.3.1: Perform a matrix calculation on the query vector of the search area and the concatenated key vector of the template and the search area, obtaining the similarity matrix between the template feature and the search area feature in the search area attention branch;

[0044] Step 2.3.2: For perform an activation function operation, and then based on and the concatenated template and the key vector value vector of the search area perform matrix calculations to obtain the most important target information in the search area features . Among them, the activation function used in the method of the present invention is softmax.

[0045] Step 3: Loop and execute Step 2 eleven times. Before each execution, use the output by the template attention branch obtained from the previous execution and the output by the search area attention branch respectively replace the initial inputs and as the inputs of the search area-guided attention mechanism for this execution.

[0046] Step 4: After looping and executing Step 2 eleven times, the search area-guided attention mechanism obtains and outputs the final features and that are most relevant to the current target state.

[0047] Step 5: The regression network predicts the target position based on the output in Step 4.

[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0049] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementing in the flow Figure 1 one process or multiple processes and / or blocksFigure 1 means for the functions specified in one or more blocks.

[0050] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0051] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one Figure 1 or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0052] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.

Claims

1. A visual target tracking method that dynamically adapts to target state changes, characterized in that: include: Step 1: Get a template image from a given video sequence and search area image ; Step 2: Dynamic feature extraction network receives template image and search area image As input, two images are mapped into two feature embeddings using the token conversion module and , and As input to the attention mechanism guided by the search region; Use the search area guided attention mechanism for dynamic feature extraction, so that the template attention branch outputs the template information most relevant to the target state in the template feature , the search area attention branch outputs the most important target information in the search area features ; Step 3: Loop through step 2 11 times. Before each execution, use the output of the template attention branch. And the output of the search area attention branch Replace the initial input and , as the input of the attention mechanism guided by the search region; Step 4: After looping through step 2 11 times, the search region-guided attention mechanism obtains and outputs the final features most relevant to the current target state. and ; Step 5: The regression network is based on the output of step 4 , predict the target location.

2. The visual target tracking method that dynamically adapts to target state changes according to claim 1 is characterized in that: The interaction process between the template attention branch and the search area attention branch in the search area guided attention mechanism in step 2 is as follows: Step 2.1: Follow and The spatial dimensions of are expanded and concatenated, using a fully connected layer mapping and , get the query vector representing the template features , key vector Sum value vector and a query vector representing the characteristics of the search area , key vector Sum value vector ; Step 2.2: The template attention branch receives the query vector of the search region with the current target state and the query vector of the template , key vector Sum value vector As input; Step 2.3: The search region attention branch receives the query vector of the search region with the current target state. and the concatenated template and search region key vector Sum value vector As input, represents the concatenation of two vectors along the spatial dimension, represents any vector.

3. The visual target tracking method that dynamically adapts to target state changes according to claim 2 is characterized in that: The query vector of the template in step 2.2 , key vector Sum value vector and the query vector for the search area The attention interaction process between them is as follows: Step 2.2.1: Query vector for the search area First and the template key vector Perform matrix calculation to obtain the similarity matrix between them , each row of the matrix represents the similarity between a search region feature and the entire template feature; Step 2.2.2: Template query vector and key vector Perform matrix calculation to obtain the similarity matrix between all template features , each row of the matrix represents the similarity between a template feature and the entire template feature; Step 2.2.3: Matrix Calculate the mean of each row of features and select the top row with the largest mean. Row (to filter out irrelevant attention responses caused by interference information between the template and the search area), and obtain the new attention matrix that is most relevant to the target feature in the template attention branch ,in, Set to number of rows; Step 2.2.4: Element-wise addition and , to introduce similarities related to the current target state into In the template attention branch, the final similarity matrix is ​​obtained ; Step 2.2.5: Matrix Perform an activation function operation to focus the weights on the information most relevant to the target, and then and the value vector of the template Perform matrix multiplication to calculate the matrix The value vector of the template weighted by the similarity weight in The target information in the template is obtained to obtain the template information most relevant to the target state in the template feature .

4. The visual target tracking method that dynamically adapts to target state changes according to claim 2 is characterized in that: The query vector of the search area in step 2.3 and the key vector of the concatenated template and search region Sum value vector The specific process of attention interaction between them is as follows: Step 2.3.1: Query vector for the search area and the concatenated template and search region key vector Perform matrix calculation to obtain the similarity matrix between the template features and the search area features in the search area attention branch ; Step 2.3.2: Execute the activation function operation, and then based on and the concatenated template and search region key vector value vector Perform matrix calculations to obtain the most important target information in the search area features .

5. The visual target tracking method that dynamically adapts to target state changes according to claim 2 is characterized in that: The network model initialization weights used by the search area guided attention mechanism include ConvMAE.

6. The visual target tracking method that dynamically adapts to target state changes according to claim 3 or 4, characterized in that: The activation function is softmax.

7. A visual target tracking device that dynamically adapts to changes in target state, characterized in that: include: Image acquisition module: obtains a template image from a given video sequence and search area image ; Dynamic feature extraction network: The dynamic feature extraction network receives the template image and search area image As input, two images are mapped into two feature embeddings using the token conversion module and , and As input to the attention mechanism guided by the search region; Use the search area guided attention mechanism for dynamic feature extraction, so that the template attention branch outputs the template information most relevant to the target state in the template feature , the search area attention branch outputs the most important target information in the search area features ; Loop execution module: loops through the dynamic feature extraction network 11 times. Before each execution, the output of the template attention branch is used. And the output of the search area attention branch Replace the initial input and , as the input of the attention mechanism guided by the search region; Feature acquisition module: After looping through the dynamic feature extraction network 11 times, the search region-guided attention mechanism obtains and outputs the final features most relevant to the current target state. and ; Prediction module: The regression network is based on the output of step 4 , predict the target location.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the visual target tracking method that dynamically adapts to target state changes as described in any one of claims 1 to 6 are implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the visual target tracking method that dynamically adapts to target state changes as described in any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the visual target tracking method that dynamically adapts to target state changes as described in any one of claims 1 to 6 is implemented.