Target tracking method and system based on mixed features and multi-scale fusion attention
By introducing mixed features and multi-scale fusion attention modules in video tracking, the problems of high-resolution video computing burden and noise interference are solved, and the accuracy and efficiency of video tracking are improved.
Patent Information
- Application Number
- CN202510814640.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Existing video tracking technologies are burdened by high-resolution videos and are susceptible to noise interference. The self-attention mechanism ignores local information, resulting in limited tracking performance.
Using a hybrid feature module and a multi-scale fusion module based on Transformer network, the global and local features of the image are extracted through the depth-separable convolution and dynamic attention weight fusion mechanism, and the feature representation is optimized by selecting attention mechanism and multi-scale fusion.
It enhances image detail capture capability, reduces computational redundancy, enriches context information, optimizes feature representation and adapts to complex scenarios, and improves the accuracy of tracking models.
Smart Images

Figure CN120339649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to an object tracking method and system based on hybrid features and multi-scale fusion attention. Background Art
[0002] Video tracking is an important research direction in the field of computer vision, and it plays a key role in many practical application fields such as autonomous driving, video surveillance, traffic management, and high-speed photography. The core task of video tracking is to accurately predict and locate a target object in subsequent frames after determining the target object in the first frame of a video sequence, so as to achieve continuous tracking. With the continuous improvement of the practicality and real-time performance of video tracking technology, its application in daily life is becoming increasingly widespread, and its research value is becoming more and more prominent. However, due to the existence of complex factors such as object deformation, rapid movement, and occlusion, video tracking is still a challenging task.
[0003] In recent years, the introduction of the Transformer architecture has brought significant progress to the field of computer vision. Through its self-attention mechanism, the Transformer can effectively explore the correlations between consecutive frames, thereby obtaining rich context information and achieving excellent tracking performance. However, a characteristic of the self-attention mechanism is that it needs to comprehensively process all input features. With the continuous increase in video and image resolution, this not only brings a greater computational burden but also may introduce additional noise interference. In addition, the self-attention mechanism mainly focuses on global information and ignores the importance of local information. Summary of the Invention
[0004] In view of the above situation, the main purpose of the present invention is to propose an object tracking method and system based on hybrid features and multi-scale fusion attention to solve the above technical problems.
[0005] The present invention proposes an object tracking method based on hybrid features and multi-scale fusion attention, and the method includes the following steps: Step 1, construct a hybrid feature module based on the Transformer network framework, construct a multi-scale fusion module based on the depthwise separable convolution and dynamic attention weight fusion mechanism, and the hybrid feature module and the multi-scale fusion module constitute a tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a secondary-stage hybrid feature module; Step 2: Input the template image and the search image into the main-stage hybrid feature module. Use the selective attention mechanism to extract the global features of the template image and the search image, use convolutional operations to extract the local features of the template image and the search image, and fuse the global features and the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module; Step 3: Input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module; Step 4: Input the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the secondary-stage hybrid feature module for interactive fusion processing to obtain the output of the secondary-stage hybrid feature module; Calculate the classification loss based on the output of the secondary-stage hybrid feature module, and use the classification loss to optimize the tracking model to obtain the optimized tracking model; Step 5: Pre-train the optimized tracking model using a large-scale dataset and adjust the model parameters to obtain the tracking model with adjusted parameters; Step 6: Input the template image and the search image into the tracking model with adjusted parameters, and repeat Steps 2 to 4 iteratively. After the iteration is completed, obtain the final output of the secondary-stage hybrid feature module; Input the final output of the secondary-stage hybrid feature module into the prediction head to obtain the tracking result.
[0006] The present invention also proposes an object tracking system based on hybrid features and multi-scale fusion attention. The system includes: A construction module for: Construct a hybrid feature module based on the Transformer network framework, and construct a multi-scale fusion module based on the depthwise separable convolution and the dynamic attention weight fusion mechanism. The hybrid feature module and the multi-scale fusion module form a tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a secondary-stage hybrid feature module; An extraction module for: Input the template image and the search image into the main-stage hybrid feature module. Use the selective attention mechanism to extract the global features of the template image and the search image, use convolutional operations to extract the local features of the template image and the search image, and fuse the global features and the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module; A calculation module for: Input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module; A learning module, configured to: Input the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module; Calculate the classification loss based on the output of the sub-stage hybrid feature module, and optimize the tracking model using the classification loss to obtain an optimized tracking model; A pre-training module, configured to: Pre-train the optimized tracking model using a large-scale dataset and adjust the model parameters to obtain a tracking model with adjusted parameters; A tracking module, configured to: Input the template image and the search image into the tracking model with adjusted parameters, and repeat the extraction module, the calculation module, and the learning module in an iterative manner as inputs. After reaching the preset number of iterations, obtain the output of the final sub-stage hybrid feature module; Input the output of the final sub-stage hybrid feature module into the prediction head to obtain the tracking result.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention extracts and fuses the global features and local features of the input image through the invented hybrid feature module, enhancing the network's ability to capture image details, enriching the context information while retaining local details; 2. The present invention uses the selective attention in the invented hybrid feature module to filter out elements with low similarity scores by setting a threshold, retaining elements with high scores, reducing the interference of irrelevant information, reducing computational redundancy, which is more obvious on large-scale datasets; 3. The present invention provides features of different scales for the input image through the invented multi-scale fusion module, making the generated feature representation more rich and comprehensive, optimizing the receptive field, and enhancing the feature representation ability and adapting to complex scenarios at the same time.
[0008] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of the steps of the object tracking method based on hybrid features and multi-scale fusion attention proposed by the present invention.
[0010] Figure 2This is the structural diagram of the object tracking framework for the object tracking method based on hybrid features and multi-scale fusion attention proposed by the present invention.
[0011] Figure 3 This is the schematic diagram of the hybrid feature module for the object tracking method based on hybrid features and multi-scale fusion attention proposed by the present invention.
[0012] Figure 4 This is the schematic diagram of the multi-scale fusion module for the object tracking method based on hybrid features and multi-scale fusion attention proposed by the present invention.
[0013] Figure 5 This is the structural diagram of the object tracking system based on hybrid features and multi-scale fusion attention proposed by the present invention. Detailed implementation manners
[0014] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0015] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed as some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0016] Please refer to Figure 1 , the embodiments of the present invention propose an object tracking method based on hybrid features and multi-scale fusion attention. The method includes the following steps: Step 1: Construct a hybrid feature module based on the Transformer network framework, and construct a multi-scale fusion module based on the depthwise separable convolution and dynamic attention weight fusion mechanism. The hybrid feature module and the multi-scale fusion module constitute a tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a secondary-stage hybrid feature module.
[0017] Step 2: Input the template image and the search image into the main-stage hybrid feature module, use the selective attention mechanism to extract the global features of the template image and the search image, use convolutional operations to extract the local features of the template image and the search image, and fuse the global features of the template image and the search image with the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module.
[0018] Please refer to Figure 2 and Figure 3, in step 2, the template image and the search image are input into the main-stage hybrid feature module. The global features of the template image and the search image are extracted by using the selective attention mechanism, and the local features of the template image and the search image are extracted by using convolution operations. Then, the global features of the template image and the search image are fused with the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module, which specifically includes the following steps: S101. Input the image into the main-stage hybrid feature module, and perform linear projection processing and dot product calculation in sequence to obtain the attention score matrix; S102. Preset an attention score matrix threshold, and screen the attention score matrix by using the attention score matrix threshold to obtain the selected attention score matrix; S103. Perform weighted calculation on the selected attention score matrix by using an activation function to obtain the globally feature after selection and weighting in the main-stage hybrid feature module; S104. Extract local features from the image through convolution operations and activation functions to obtain the local features extracted in the main-stage hybrid feature module; S105. After concatenating and linearly processing the globally feature after selection and weighting in the main-stage hybrid feature module and the local features extracted in the main-stage hybrid feature module in sequence, obtain the output of the main-stage hybrid feature module; S106. Repeat S101 to S105 with the template image and the search image as inputs to obtain the outputs of the main-stage hybrid feature modules of the template image and the search image.
[0019] Input the template image and the search image into the main-stage hybrid feature module respectively, and perform linear projection processing and dot product calculation in sequence to obtain the attention score matrix. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the attention score matrix, represents the similarity between the query and the corresponding key, represents the query matrix, represents the key matrix, represents transpose, represents dimension; It should be noted that the currently input feature vector, is the feature vector that is matched with
[0020] In the step of screening the attention score matrix by using the attention score matrix threshold to obtain the selected attention score matrix, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the result after selection of the th row and th column in the attention score matrix, represents the set threshold; In the step of using the activation function to perform weighted calculation on the selected attention score matrix to obtain the globally weighted feature in the main stage hybrid feature module, the relational expression in the corresponding process is as follows: ; Among them, represents the globally weighted feature in the main stage hybrid feature module, represents the first activation function, represents the attention score matrix after selection by the set threshold, represents the value matrix; In the step of performing local feature extraction on the template image and the search image through convolution operation and activation function to obtain the local feature extracted in the main stage hybrid feature module, the relational expression in the corresponding process is as follows: ; Among them, represents the first local feature extracted in the main stage hybrid feature module, represents the second local feature extracted in the main stage hybrid feature module, represents the convolution operation with a convolution kernel size of 3x3, represents the second activation function, represents the input feature; In the step of sequentially performing splicing processing and linear layer processing on the globally weighted feature in the main stage hybrid feature module and the local feature extracted in the main stage hybrid feature module to obtain the output of the main stage hybrid feature module, the relational expression in the corresponding process is as follows: ; Among them, represents the output of the main stage hybrid feature module, represents the linear layer, represents the splicing operation.
[0021] Furthermore, the main stage hybrid feature module extracts global features through selective attention, highlights important features and suppresses unimportant features, improves the model's perception ability of key information, reduces computational redundancy, and the fusion of global features and local features enables the model to have a richer feature representation. This type of fusion method enhances global context information while retaining local details.
[0022] Step 3: Input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module.
[0023] Please refer to Figure 4 , in Step 3, input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module, which specifically includes the following steps: Input the output of the main-stage hybrid feature module into the multi-scale fusion module to perform feature extraction using depthwise separable convolution to obtain two extracted features of different scales; Concatenate the two extracted features of different scales to obtain the concatenated feature; Send the concatenated feature into the channel feature selection branch, and perform channel average pooling and channel max pooling respectively to obtain the channel average pooling feature and the channel max pooling feature. Then, perform concatenation processing and convolution processing on the channel average pooling feature and the channel max pooling feature in sequence to obtain two channel attention features; Activate the two channel attention features using an activation function to obtain two channel masks; Send the concatenated feature into the spatial feature selection branch, and perform spatial global average pooling, fully connected layer processing, and activation processing in sequence to obtain two spatial masks; Multiply the two channel masks and the two spatial masks to obtain two dynamically adjustable weights; Perform weighted fusion processing on the two dynamically adjustable weights and the two extracted features of different scales respectively to obtain a weighted feature map, and perform convolution processing on the weighted feature map to obtain the output of the multi-scale fusion module.
[0024] Input the output of the main-stage hybrid feature module into the multi-scale fusion module to perform feature extraction using depthwise separable convolution to obtain two extracted features of different scales. The corresponding relationship existing in the process is as follows: ; Among them, represents the first extracted feature of different scales, represents the second extracted feature of different scales, represents a convolution operation with a convolution kernel size of 1x1, represents a depthwise separable convolution operation with a convolution kernel size of 5x5, represents a depthwise separable convolution operation with a convolution kernel size of 3x3, represents the feature input from the main-stage hybrid feature module to the multi-scale fusion module; In the step of splicing two extracted features of different scales to obtain the spliced features, the relational expressions in the corresponding process are as follows: ; Among them, represents the spliced features; In the step of sending the spliced features into the channel feature selection branch, performing channel average pooling and channel max pooling respectively to obtain channel average pooling features and channel max pooling features, and then performing splicing and convolution operations on the channel average pooling features and channel max pooling features in sequence to obtain two channel attention features, the relational expressions in the corresponding process are as follows: ; Among them, represents the channel attention feature obtained after the convolution operation on the spliced features, represents the channel average pooling operation, represents the channel max pooling operation; In the step of activating two channel attention features using an activation function to obtain two channel masks, the relational expressions in the corresponding process are as follows: ; Among them, represents the th channel mask, represents the channel mask index, represents the second activation function, represents the th channel attention feature, represents the channel attention feature index.
[0025] Sending the spliced features into the spatial feature selection branch, performing spatial global average pooling, fully connected layer processing and activation processing in sequence to obtain two spatial masks, the relational expressions in the corresponding process are as follows: ; Among them, represents the spatial attention feature generated after the fully connected layer processing, represents the fully connected layer, represents the spatial global average pooling, represents the th spatial mask, represents the spatial mask index, represents the th spatial attention feature, represents the spatial attention feature index; In the step of multiplying two channel masks and two spatial masks to obtain two dynamically adjustable weights, the relational expressions in the corresponding process are as follows: ; Where, represents the first dynamically adjustable weight, represents the second dynamically adjustable weight, represents the first channel mask, represents the second channel mask, represents the first spatial mask, represents the second spatial mask; In the step of respectively performing weighted fusion processing on two dynamically adjustable weights and two extracted features of different scales to obtain a weighted feature map, and performing convolution processing on the weighted feature map to obtain the output of the multi-scale fusion module, the relational expressions in the corresponding process are as follows: ; Where, represents the output of the multi-scale fusion module.
[0026] Furthermore, the multi-scale fusion module uses depthwise separable convolutions with different-sized convolutional kernels to obtain features of different scales, obtaining context features with different receptive fields. The features of different scales are passed through a convolutional layer to integrate channel information and unify the dimensions. The features with different receptive fields are concatenated and then sent into a channel feature selection branch and a spatial feature selection branch to generate two dynamically adjustable weights; the channel feature selection branch respectively processes the concatenated features through max pooling and average pooling, concatenates the features after max pooling and average pooling, and the concatenated features pass through a convolutional layer to obtain two attention features. The two attention features are processed through a Sigmoid activation function to obtain two channel masks; the spatial feature selection branch sequentially processes the concatenated features with different receptive fields through global average pooling, a fully connected layer, and a Softmax activation function to generate two spatial masks; the generated channel masks and spatial masks are multiplied to obtain two dynamically adjustable weights. The features processed by different convolutional kernels are weighted and fused with the two weights to obtain a weighted feature map, and the weighted feature map is processed through a convolution operation to achieve feature fusion at different scales.
[0027] Step 4: Input the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module; Calculate the classification loss based on the output of the sub-stage hybrid feature module, and optimize the tracking model using the classification loss to obtain an optimized tracking model.
[0028] In step 4, the outputs of the main-stage hybrid feature module and the multi-scale fusion module are input into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module, which specifically includes the following steps: The outputs of the main-stage hybrid feature module and the multi-scale fusion module are input into the sub-stage hybrid feature module to obtain image features, and the template image sequence feature and the search image sequence feature are obtained respectively; The template image sequence feature and the search image sequence feature are concatenated to obtain the feature after concatenating the image sequences. The feature after concatenating the image sequences is used as the input to repeat steps S101 to S105 for selective attention calculation, and the globally weighted feature and the concatenated value matrix in the sub-stage hybrid feature module are obtained respectively; Convolution operations, activation processing, and convolution processing are performed on the feature after concatenating the image sequences and the concatenated value matrix respectively to obtain two local features extracted by the sub-stage hybrid feature module; The globally weighted feature in the sub-stage hybrid feature module and the two local features extracted by the sub-stage hybrid feature module are sequentially concatenated and linearly processed to obtain the output of the sub-stage hybrid feature module.
[0029] The template image sequence feature and the search image sequence feature are concatenated to obtain the feature after concatenating the image sequences. The feature after concatenating the image sequences is used as the input to repeat steps S101 to S105 for selective attention calculation, and the globally weighted feature and the concatenated value matrix in the sub-stage hybrid feature module are obtained respectively. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the feature after concatenating the image sequences, represents the template image sequence feature, represents the search image sequence feature, represents the concatenated query matrix, the query matrix of the template image, represents the query matrix of the search image, represents the concatenated key matrix, represents the key matrix of the template image, represents the key matrix of the search image, represents the concatenated value matrix, represents the value matrix of the template image, represents the value matrix of the search image, represents the attention score matrix in the sub-stage hybrid feature module, represents the similarity between the concatenated query matrix and the corresponding concatenated key matrix, Represents the attention score matrix in the sub-stage hybrid feature module The result after selecting the th row and the th column, represents the global feature after selection and weighting in the sub-stage hybrid feature module; In the step of performing convolution operations, activation processing, and convolution processing on the concatenated features of the image sequence and the concatenated value matrix respectively to obtain two local features extracted by the sub-stage hybrid feature module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the first local feature extracted by the sub-stage hybrid feature module, represents the second local feature extracted by the sub-stage hybrid feature module; In the step of sequentially performing concatenation processing and linear layer processing on the globally weighted feature in the sub-stage hybrid feature module and the two local features extracted by the sub-stage hybrid feature module to obtain the output of the sub-stage hybrid feature module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the output of the sub-stage hybrid feature module.
[0030] Furthermore, the sub-stage hybrid feature module is used to interactively fuse the template image features and search image features processed by the main-stage hybrid feature module and the multi-scale fusion module. In the sub-stage hybrid feature module, the global features and local features of the input template image and search image are simultaneously extracted through the hybrid feature module, and the extracted global features and local features are concatenated, enhancing the global context information while retaining local details. The hybrid feature module extracts global features, highlighting important features and suppressing unimportant features, improving the tracking model's perception ability of key information and reducing computational redundancy. The fusion of global features and local features enables the model to have a richer feature representation.
[0031] Step 5: Use a large-scale dataset to pre-train the optimized tracking model and adjust the model parameters to obtain a tracking model with adjusted parameters.
[0032] Furthermore, randomly select 60,000 pictures from four public datasets, namely TrackingNet, GOT-10k, LaSOT, and COCO2017, for training. The number of training epochs is 300. When training reaches the 240th epoch, reduce the learning rate to 10% of the initial value to enable the model to converge quickly.
[0033] Step 6: Input the template image and the search image into the tracking model with adjusted parameters, and repeat Steps 2 to 4 iteratively. After reaching the preset number of iterations, obtain the output of the final sub-stage hybrid feature module; Input the output of the final sub-stage hybrid feature module into the prediction head to obtain the tracking result.
[0034] Please refer to Figure 5 , the present invention also proposes an object tracking system based on hybrid features and multi-scale fusion attention. The system includes: A construction module for: Construct a hybrid feature module based on the Transformer network framework, and construct a multi-scale fusion module based on the depthwise separable convolution and dynamic attention weight fusion mechanism. The hybrid feature module and the multi-scale fusion module constitute the tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a sub-stage hybrid feature module; An extraction module for: Input the template image and the search image into the main-stage hybrid feature module, extract the global features of the template image and the search image using the selective attention mechanism, extract the local features of the template image and the search image using convolution operations, and fuse the global features of the template image and the search image with the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module; A calculation module for: Input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module; A learning module for: Input the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module; Calculate the classification loss according to the output of the sub-stage hybrid feature module, and optimize the tracking model using the classification loss to obtain the optimized tracking model; A pre-training module for: Pre-train the optimized tracking model using a large-scale dataset and adjust the model parameters to obtain the tracking model with adjusted parameters; A tracking module for: Input the template image and the search image into the tracking model with adjusted parameters, and repeat the extraction module, the calculation module, and the learning module iteratively as inputs. After reaching the preset number of iterations, obtain the output of the final sub-stage hybrid feature module; Input the output of the final sub-stage hybrid feature module into the prediction head to obtain the tracking result.
[0035] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0036] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0037] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A target tracking method based on hybrid features and multi-scale fusion attention, characterized in that, The method includes the following steps: Step 1: Construct a hybrid feature module based on the Transformer network framework, and construct a multi-scale fusion module based on depthwise separable convolution and dynamic attention weight fusion mechanism. The hybrid feature module and the multi-scale fusion module constitute a tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a sub-stage hybrid feature module; Step 2: Input the template image and the search image into the main-stage hybrid feature module, use the selective attention mechanism to extract the global features of the template image and the search image, use convolution operations to extract the local features of the template image and the search image, and fuse the global features of the template image and the search image with the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module; Step 3: Input the output of the main-stage hybrid feature module into the multi-scale fusion module, and combine the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module; Step 4: Input the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module; Calculate the classification loss according to the output of the sub-stage hybrid feature module, and use the classification loss to optimize the tracking model to obtain an optimized tracking model; Step 5: Pre-train the optimized tracking model using a large-scale dataset and adjust the model parameters to obtain a tracking model with adjusted parameters; Step 6: Input the template image and the search image into the tracking model with adjusted parameters, and repeat steps 2 to 4 iteratively. After reaching the preset number of iterations, obtain the final output of the sub-stage hybrid feature module; Input the final output of the sub-stage hybrid feature module into the prediction head to obtain the tracking result.
2. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 1, characterized in that, In step 2, when inputting the template image and the search image into the main-stage hybrid feature module, using the selective attention mechanism to extract the global features of the template image and the search image, using convolution operations to extract the local features of the template image and the search image, and fusing the global features of the template image and the search image with the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module, specifically includes the following steps: S101: Input the image into the main-stage hybrid feature module, perform linear projection processing and dot product calculation in sequence to obtain an attention score matrix; S102: Preset an attention score matrix threshold, and sieve the attention score matrix using the attention score matrix threshold to obtain a selected attention score matrix; S103: Perform weighted calculation on the selected attention score matrix using an activation function to obtain the globally feature after selection and weighting in the main-stage hybrid feature module; S104: Extract local features from the image through convolution operations and activation functions to obtain the locally features extracted in the main-stage hybrid feature module; S105. After successively performing splicing processing and linear layer processing on the globally weighted features after selection weighting in the main-stage hybrid feature module and the locally extracted features in the main-stage hybrid feature module, the output of the main-stage hybrid feature module is obtained. S106. Using the template image and the search image as inputs, repeat S101 to S105 to obtain the outputs of the main-stage hybrid feature modules of the template image and the search image.
3. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 2, wherein The template image and the search image are respectively input into the main-stage hybrid feature module, and linear projection processing and dot product calculation are successively performed to obtain an attention score matrix. The relational expressions existing in the corresponding process are as follows: ; Among them, represents the attention score matrix, represents the similarity between the query and the corresponding key, represents the query matrix, represents the key matrix, represents the transpose, represents the dimension; In the step of screening the attention score matrix using the attention score matrix threshold to obtain the selected attention score matrix, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the attention score matrix in the row after column selection, represents the set threshold; In the step of performing weighted calculation on the selected attention score matrix using an activation function to obtain the globally weighted features after selection weighting in the main-stage hybrid feature module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the global feature after selection and weighting in the main stage hybrid feature module, represents the first activation function, represents the attention score matrix after selection by the set threshold, represents the value matrix; In the step of extracting local features from the template image and the search image through convolution operations and activation functions to obtain the locally extracted features in the main-stage hybrid feature module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the first extracted local feature in the main-stage hybrid feature module, represents the second extracted local feature in the main-stage hybrid feature module, represents a convolution operation with a convolution kernel size of 3x3, represents the second activation function, represents the input feature; In the step of successively performing splicing processing and linear layer processing on the globally weighted features after selection weighting in the main-stage hybrid feature module and the locally extracted features in the main-stage hybrid feature module to obtain the output of the main-stage hybrid feature module, the relational expressions existing in the corresponding process are as follows: ; Among them, represents the output of the main stage hybrid feature module, represents the linear layer, represents the concatenation operation.
4. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 3, wherein In step 3, the output of the main-stage hybrid feature module is input into the multi-scale fusion module, and combined with the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch, the output of the multi-scale fusion module is obtained. The specific steps are as follows: The output of the main-stage hybrid feature module is input into the multi-scale fusion module for feature extraction using depthwise separable convolution to obtain two extracted features of different scales. The two extracted features of different scales are spliced to obtain the spliced feature. The spliced feature is sent to the channel feature selection branch, and channel average pooling processing and channel max pooling processing are respectively performed to obtain a channel average pooling feature and a channel max pooling feature. The channel average pooling feature and the channel max pooling feature are successively subjected to splicing processing and convolution processing to obtain two channel attention features. The two channel attention features are activated using an activation function to obtain two channel masks. The spliced feature is sent to the spatial feature selection branch, and spatial global average pooling processing, fully connected layer processing, and activation processing are successively performed to obtain two spatial masks. The two channel masks and the two spatial masks are multiplied to obtain two dynamically adjustable weights. The two dynamically adjustable weights are respectively subjected to weighted fusion processing with the two extracted features of different scales to obtain a weighted feature map, and the weighted feature map is subjected to convolution processing to obtain the output of the multi-scale fusion module.
5. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 4, wherein The output of the main-stage hybrid feature module is input into the multi-scale fusion module for feature extraction using depthwise separable convolution to obtain two extracted features of different scales. The relational expressions for the corresponding process are as follows: ; Among them, represents the first extracted different-scale feature, represents the second extracted different-scale feature, represents a convolution operation with a convolution kernel size of 1x1, represents a depthwise separable convolution operation with a convolution kernel size of 5x5, represents a depthwise separable convolution operation with a convolution kernel size of 3x3, represents the feature input from the main-stage mixed feature module to the multi-scale fusion module; In the step of concatenating the two extracted features of different scales to obtain the concatenated feature, the relational expressions for the corresponding process are as follows: ; Among them, represents the spliced feature; In the step of feeding the concatenated feature into the channel feature selection branch to perform channel average pooling and channel max pooling respectively to obtain the channel average pooling feature and the channel max pooling feature, and then performing concatenation processing and convolution processing on the channel average pooling feature and the channel max pooling feature in sequence to obtain two channel attention features, the relational expressions for the corresponding process are as follows: ; Among them, represents the channel attention feature obtained after the convolutional operation on the spliced features, represents the channel average pooling operation, represents the channel maximum pooling operation; In the step of activating the two channel attention features using an activation function to obtain two channel masks, the relational expressions for the corresponding process are as follows: ; Among them, represents the th channel mask, represents the channel mask index, represents the second activation function, represents the th channel attention feature, represents the channel attention feature index.
6. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 5, characterized in that The concatenated feature is fed into the spatial feature selection branch to perform spatial global average pooling, fully connected layer processing, and activation processing in sequence to obtain two spatial masks. The relational expressions for the corresponding process are as follows: ; Among them, represents the spatial attention feature generated after being processed by the fully connected layer, represents the fully connected layer, represents the spatial global average pooling process, represents the th spatial mask, represents the spatial mask index, represents the th spatial attention feature, represents the spatial attention feature index; In the step of multiplying the two channel masks and the two spatial masks to obtain two dynamically adjustable weights, the relational expressions for the corresponding process are as follows: ; Among them, represents the first dynamically adjustable weight, represents the second dynamically adjustable weight, represents the first channel mask, represents the second channel mask, represents the first spatial mask, represents the second spatial mask; In the step of performing weighted fusion processing on the two dynamically adjustable weights and the two extracted features of different scales respectively to obtain a weighted feature map, and then performing convolution processing on the weighted feature map to obtain the output of the multi-scale fusion module, the relational expressions for the corresponding process are as follows: ; Among them, represents the output of the multi-scale fusion module.
7. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 6, characterized in that, In step 4, the output of the main-stage hybrid feature module and the output of the multi-scale fusion module are input into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module. Specifically, it includes the following steps: The output of the main-stage hybrid feature module and the output of the multi-scale fusion module are input into the sub-stage hybrid feature module to obtain image features, and the template image sequence feature and the search image sequence feature are obtained respectively; The template image sequence feature and the search image sequence feature are concatenated to obtain the feature after image sequence concatenation. The feature after image sequence concatenation is used as the input to repeat steps S101 to S105 to calculate selective attention, and the globally weighted feature and the concatenated value matrix in the sub-stage hybrid feature module are obtained respectively; Convolution operations, activation processing, and convolution processing are performed on the feature after image sequence concatenation and the concatenated value matrix respectively to obtain two local features extracted by the sub-stage hybrid feature module; The globally weighted feature in the sub-stage hybrid feature module and the two local features extracted by the sub-stage hybrid feature module are concatenated and linearly processed in sequence to obtain the output of the sub-stage hybrid feature module.
8. The object tracking method based on hybrid features and multi-scale fusion attention according to claim 7, characterized in that Perform splicing processing on the features of the template image sequence and the features of the search image sequence to obtain the features after concatenating the image sequences. Use the features after concatenating the image sequences as the input to repeat steps S101 to S105 to calculate selective attention, and respectively obtain the globally weighted features and the spliced value matrix in the sub-stage hybrid feature module. The relational expressions in the corresponding process are as follows: ; Among them, represents the feature after concatenating the image sequences, represents the feature of the template image sequence, represents the feature of the search image sequence, represents the concatenated query matrix, the query matrix of the template image, represents the query matrix of the search image, represents the concatenated key matrix, represents the key matrix of the template image, represents the key matrix of the search image, represents the concatenated value matrix, represents the value matrix of the template image, represents the value matrix of the search image, represents the attention score matrix in the sub-stage hybrid feature module, represents the similarity between the concatenated query matrix and the corresponding concatenated key matrix, represents the attention score matrix in the sub-stage hybrid feature module The row and the column selection result, represents the globally weighted feature after selection in the sub-stage hybrid feature module, represents the attention score matrix after being selected by setting a threshold in the sub-stage hybrid feature module; In the step of performing convolution operations, activation processing, and convolution processing on the features after concatenating the image sequences and the spliced value matrix respectively to obtain two local features extracted by the sub-stage hybrid feature module, the relational expressions in the corresponding process are as follows: ; Among them, represents the first local feature extracted by the sub-stage hybrid feature module, represents the second local feature extracted by the sub-stage hybrid feature module; In the step of performing splicing processing and linear layer processing on the globally weighted features in the sub-stage hybrid feature module and the two local features extracted by the sub-stage hybrid feature module in sequence to obtain the output of the sub-stage hybrid feature module, the relational expressions in the corresponding process are as follows: ; Among them, represents the output of the sub-stage hybrid feature module.
9. An object tracking system based on hybrid features and multi-scale fusion attention, characterized in that, The system applies the object tracking method based on hybrid features and multi-scale fusion attention as described in any one of the above claims 1 to 8. The system includes: A construction module for: Constructing a hybrid feature module based on the Transformer network framework, and constructing a multi-scale fusion module based on the depthwise separable convolution and the dynamic attention weight fusion mechanism. The hybrid feature module and the multi-scale fusion module constitute a tracking model; Among them, the hybrid feature module includes a main-stage hybrid feature module and a sub-stage hybrid feature module; An extraction module for: Inputting the template image and the search image into the main-stage hybrid feature module, using the selective attention mechanism to extract the global features of the template image and the search image, using convolution operations to extract the local features of the template image and the search image, and fusing the global features and the local features of the template image and the search image to obtain the output of the main-stage hybrid feature module; A calculation module for: Inputting the output of the main-stage hybrid feature module into the multi-scale fusion module, and combining the dynamic weights generated by the channel feature selection branch and the spatial feature selection branch to obtain the output of the multi-scale fusion module; A learning module for: Inputting the output of the main-stage hybrid feature module and the output of the multi-scale fusion module into the sub-stage hybrid feature module for interactive fusion processing to obtain the output of the sub-stage hybrid feature module; Calculating the classification loss according to the output of the sub-stage hybrid feature module, and optimizing the tracking model using the classification loss to obtain an optimized tracking model; A pre-training module for: Pre-training the optimized tracking model using a large-scale dataset and adjusting the model parameters to obtain a tracking model with adjusted parameters; A tracking module for: Inputting the template image and the search image into the tracking model with adjusted parameters, and iteratively repeating the extraction module, the calculation module, and the learning module as inputs. After reaching the preset number of iterations, obtain the final output of the sub-stage hybrid feature module; Inputting the final output of the sub-stage hybrid feature module into the prediction head to obtain the tracking result.
Citation Information
Patent Citations
Transform-based high-performance target tracking method
CN118982561A
Target tracking method and system based on grouping attention feature extraction network
CN119273941A
Method and system for tracking object by aggregation network based on hybrid convolution and self-attention
US20240104772A1