Target tracking method and system based on spatial channel summation attention
By introducing a dual-branch feature extraction network and feature fusion network of spatial channel summation attention in visual tracking, the problem of ignoring global context information and local information in the prior art is solved, and more accurate target tracking is achieved, and complex appearance changes can be dealt with.
Patent Information
- Application Number
- CN202510171644.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Existing visual tracking algorithms ignore the close connection between global context information and local information, resulting in information loss, and the large use of self-attention leads to redundant calculations, making it difficult to deal with complex appearance changes.
A target tracking method based on spatial channel summation attention is proposed. By constructing a spatial channel summation attention module and feature integration module, a dual-branch feature extraction network is built, and a variant self-attention mechanism is used to build a feature fusion network to realize information interaction between the spatial domain and the channel domain, and obtain global context information.
Through the dual-branch feature extraction network and feature fusion network of spatial channel summation attention, the global context information and local information of the image are effectively captured, and the accuracy of target tracking is improved, and the difficulties such as occlusion, rapid movement and complex background are better met.
Smart Images

Figure CN119648749B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and image processing technology, and in particular to a target tracking method and system based on spatial channel summation attention. Background Art
[0002] In recent years, visual tracking has been a research hotspot in the field of computer vision and is widely used in related fields such as unmanned driving, surveillance security, and intelligent robots. The purpose of visual tracking is to predict the bounding box position of the target of interest in subsequent frames when the target bounding box is initialized in the first frame. At present, visual tracking has made significant progress. However, it is still a difficult visual task because there are many uncertain interference factors in the real environment, such as large-area occlusion, illumination changes, and severe deformation, which will have an adverse effect on the tracking effect. With the development of deep learning and machine learning, related fields have flourished, and visual tracking has also been greatly promoted.
[0003] In video sequences, there are a large amount of feature correlation and contextual information, both of which are crucial for visual tracking. Therefore, the key to improving visual tracking is how to effectively find and utilize relevant information. In the classic twin network tracker, the correlation module plays an important role in integrating the object feature information from the template and the search area. However, the correlation module usually uses a convolutional neural network to calculate the extracted features to obtain a response map. This operation is a linear matching process that only considers a small range of local feature information and ignores a large range of global feature information. Therefore, the tracker is likely to be unable to effectively develop the complex nonlinear relationship between the template and the search area, resulting in information loss. In the face of long-term tracking and some difficulties such as occlusion and deformation, it often leads to tracking drift or failure.
[0004] Transformer has an advantage over recurrent neural networks (RNNs) in processing continuous tasks due to its excellent parallel computing capabilities and storage mechanisms. Therefore, the Transformer architecture is becoming more and more popular in the field of natural language processing. In fact, Transformer can traverse all units in a sequence with the help of an attention-based encoder and decoder to learn the dependencies between them. This feature makes Transformer naturally good at obtaining global information in a sequence. At the same time, the Transformer architecture is gradually being used for non-continuous tasks, and the application areas of Transformer have been widely expanded. The pioneering work of Vision Transformer (ViT) successfully introduced the Transformer architecture into the field of computer vision for the first time, and introduced a self-attention module to explicitly model the long-range dependencies between image blocks, overcoming the inherent limitations of the local receptive field in convolution, thereby improving the performance of various tasks.
[0005] However, current visual tracking algorithms ignore the close connection between global contextual information and local information, resulting in the loss of a large amount of local information. In addition, the extensive use of self-attention, matrix multiplication, and complex operations such as Softmax also lead to a lot of redundant calculations, which makes it difficult to deal with problems caused by complex appearance changes. Summary of the invention
[0006] In view of the above situation, the main purpose of the present invention is to propose a target tracking method and system based on spatial channel summation attention to solve the above technical problems.
[0007] The present invention proposes a target tracking method based on spatial channel summation attention, the method comprising the following steps:
[0008] Step 1: Based on the spatial attention mechanism and the channel attention mechanism, a spatial channel sum attention module is constructed. Based on the spatial channel sum attention module and the feature integration module, a dual-branch feature extraction network is constructed. Based on the variant self-attention mechanism, a feature fusion network is constructed. The dual-branch feature extraction network, the feature fusion network and the head network constitute a target tracking model.
[0009] Step 2: Pre-train the dual-branch feature extraction network and the feature fusion network using a large-scale data set to obtain a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network;
[0010] Step 3: Input the image features into the pre-trained dual-branch feature extraction network, and use the feature integration module to fuse the image features to obtain fused image features;
[0011] Perform residual connection on the image features and fused image features, and perform layer normalization. Then, send the layer normalization results to the spatial channel sum attention module for information interaction between the spatial domain and the channel domain to obtain the attention output.
[0012] The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map;
[0013] Step 4: preprocess the template image of the first frame and the search area image of the subsequent frames, perform image block embedding operations on the template image and the search area image respectively, and encode the embedding position of each image block to obtain template features and search area features;
[0014] Input the template features and search area features into the two branches of the pre-trained dual-branch feature extraction network respectively, and repeat step 3 several times in an iterative manner to obtain the final template features and intermediate search features;
[0015] Step 5: Input the final template features and the intermediate search features into the feature fusion network for fusion to obtain a two-dimensional feature map, and input the two-dimensional feature map into the head network to obtain the target tracking frame;
[0016] Step 6: Repeat steps 3 to 5 to train the target tracking model using the training set as input to obtain a trained target tracking model;
[0017] Step 7: Use the trained target tracking model to track the target.
[0018] The present invention also proposes a target tracking system based on spatial channel summation and attention, wherein the system applies the target tracking method based on spatial channel summation and attention as described above, and the system comprises:
[0019] Building blocks for:
[0020] A spatial channel sum attention module is built based on the spatial attention mechanism and the channel attention mechanism. A dual-branch feature extraction network is built based on the spatial channel sum attention module and the feature integration module. A feature fusion network is built based on the variant self-attention mechanism. The dual-branch feature extraction network, the feature fusion network and the head network constitute the target tracking model.
[0021] Training modules for:
[0022] Using a large-scale data set to pre-train a dual-branch feature extraction network and a feature fusion network, a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network are obtained;
[0023] Extraction module 1, used for:
[0024] The image features are input into the pre-trained dual-branch feature extraction network, and the image features are fused using the feature integration module to obtain fused image features;
[0025] Perform residual connection on the image features and fused image features, and perform layer normalization. Then, send the layer normalization results to the spatial channel sum attention module for information interaction between the spatial domain and the channel domain to obtain the attention output.
[0026] The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map;
[0027] Extraction module 2, used for:
[0028] Preprocessing the template image of the first frame and the search area image of the subsequent frames, performing image block embedding operations on the template image and the search area image respectively, and encoding the embedding position of each image block to obtain template features and search area features;
[0029] The template features and search area features are respectively input into the two branches of the pre-trained dual-branch feature extraction network, and the operation of a part of the extraction module is repeated several times in an iterative manner to obtain the final template features and intermediate search features;
[0030] Compute module for:
[0031] The final template features and the intermediate search features are input into the feature fusion network for fusion to obtain a two-dimensional feature map, and the two-dimensional feature map is input into the head network to obtain the target tracking frame;
[0032] Learning modules for:
[0033] The target tracking model is trained by using the training set as input to obtain a trained target tracking model;
[0034] Tracking module for:
[0035] Use the trained target tracking model to perform target tracking.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] 1. By selecting spatial attention and channel attention to construct a spatial channel sum attention module, the spatial domain information and channel domain information of image features are effectively interacted to obtain a large amount of structured spatial information and channel information, thereby obtaining global context information of different dimensions of the image.
[0038] 2. Construct a feature fusion network based on variant self-attention; concatenate the features obtained by the feature extraction network into a feature sequence, and pass it into the feature fusion network constructed by variant self-attention for feature fusion to capture more comprehensive image features.
[0039] 3. The present invention uses a dual-branch feature extraction network based on the spatial channel summation attention module to extract image features. During the extraction process, the spatial attention mechanism and the channel attention mechanism are processed to realize information interaction between the spatial domain and the channel domain, obtain a large amount of structured spatial information and channel information, obtain the final template features and intermediate search features, and then fuse the final template features and the intermediate search features and send them to the head prediction network. The maximum response position of the tracking target in the search area can be obtained, thereby tracking the target. The present invention makes full use of the advantages of spatial channel summation attention, so that the tracker can well cope with the difficulties such as target occlusion, rapid movement, and complex background that occur during the tracking process, and can achieve more accurate tracking.
[0040] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description or learned through embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flow chart of the target tracking method based on spatial channel summation attention proposed by the present invention;
[0042] Figure 2 This is a structural diagram of the target tracking framework based on spatial channel summation attention proposed by the present invention;
[0043] Figure 3 This is a structural diagram of the spatial channel summation attention module proposed in the present invention;
[0044] Figure 4 This is a schematic diagram of the structure of the target tracking system based on spatial channel summation attention proposed in the present invention. DETAILED DESCRIPTION
[0045] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0046] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.
[0047] See also Figure 1 and Figure 2 , this embodiment provides a target tracking method based on spatial channel summation attention, the method comprising the following steps:
[0048] Step 1: Based on the spatial attention mechanism and the channel attention mechanism, a spatial channel sum attention module is constructed. Based on the spatial channel sum attention module and the feature integration module, a dual-branch feature extraction network is constructed. Based on the variant self-attention mechanism, a feature fusion network is constructed. The dual-branch feature extraction network, the feature fusion network and the head network constitute a target tracking model.
[0049] Among them, the dual-branch feature extraction network includes a template branch feature extraction network for processing template images and a search branch feature extraction network for processing search area images. The template branch feature extraction network is set to 9 layers, and the search branch feature extraction network is set to 6 layers.
[0050] Among them, the feature integration module consists of a first 1×1 convolution, a normalization layer, a GeLU activated depth-separable convolution layer, and a second 1×1 convolution, which are arranged in sequence.
[0051] Step 2: Pre-train the dual-branch feature extraction network and the feature fusion network using a large-scale data set to obtain a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network;
[0052] Step 3: Input the image features into the pre-trained dual-branch feature extraction network, and use the feature integration module to fuse the image features to obtain fused image features;
[0053] Perform residual connection on the image features and fused image features, and perform layer normalization. Then, send the layer normalization results to the spatial channel sum attention module for information interaction between the spatial domain and the channel domain to obtain the attention output.
[0054] The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map;
[0055] Among them, the template image features and search area image features obtained after image preprocessing enter the nine-layer feature extraction network and the six-layer feature extraction network respectively. First, enter the feature integration module. Before feature extraction, each layer of the feature integration module first fuses the input data, and helps the model learn richer and more complex feature representations through dimensionality increase and dimensionality reduction operations, as well as deep separable convolutions, to achieve the effect of feature extraction and fusion. At the same time, the performance and stability of the model are improved without increasing too many parameters. The normalization operation is used to improve the convergence speed and performance of the model. The output features after the feature integration module and the image features before input are connected by residual connection, and the input image features are directly passed to the subsequent layers of the feature integration module, which helps the network capture deeper contextual information and accelerates convergence. The layer normalization operation is used to stabilize the input of each layer, thereby slowing down the problem of gradient disappearance or gradient explosion. Then it is passed to the spatial channel sum attention module, and the spatial channel sum attention is used to extract features, thereby realizing information interaction between the spatial domain and the channel domain.
[0056] As a further preferred embodiment of the present invention, the layer normalization result is sent to the spatial channel sum attention module to perform information interaction between the spatial domain and the channel domain, and obtaining the attention output specifically includes the following steps:
[0057] Normalizing the layer results in generating a first query, a first key, and a first value;
[0058] Using spatial attention for the first query and the first key, respectively, to assign different weights to different spatial positions of the input feature map, and obtain the spatial attention outputs of the first query and the first key;
[0059] The spatial attention outputs of the first query and the first key are respectively assigned different weights to different channels of the input feature map using channel attention to obtain the channel attention outputs of the first query and the first key;
[0060] Then add the channel attention outputs of the first query and the first key, perform a linear transformation to integrate the context information, and then perform a dot product operation on the linear transformation result and the first value to obtain the attention output.
[0061] Among them, the layer normalization result is sent to the spatial channel sum attention module for information interaction in the spatial domain and channel domain. The process of obtaining the corresponding attention output has the following relationship:
[0062] ;
[0063] ;
[0064] in, Represent the first query, first key, and first value respectively. , , , Indicates the number of image blocks, Represents the feature dimension of each image block; Respectively represent the weight matrices corresponding to the first query, the first key, and the first value; Represents the linear transformation operation used to integrate context information, which can be implemented by using depthwise separable grouped convolution. By extracting features independently on each channel, the context score is extracted. While reducing the number of parameters and the amount of computation, the depthwise separable convolution maintains the independence and diversity of features, which helps to improve computational efficiency. Represents the context mapping function, which includes basic information interaction, composed of channel attention and spatial attention composition; Represents the spatial attention calculation operation, which is used to process the data of each channel independently, assign different weights to different spatial positions of the input feature map, and emphasize important areas; Represents the channel attention calculation operation, which is used to assign different weights to different channels of the input feature map to emphasize important features; represents the attention output, Embedded representation of the input sequence of the spatial channel sum attention module, i.e., layer normalization result.
[0065] Finally, the attention output obtained after the spatial channel summation attention module is subjected to layer normalization and multi-layer perceptron operation. The layer normalization operation is used to stabilize the input of each layer so that the activation value remains within a stable range to prevent the problem of gradient explosion or gradient disappearance in the subsequent multi-layer perceptron operation; the multi-layer perceptron operation learns and extracts more complex nonlinear mapping relationships in the image through nonlinear transformations of multiple hidden layers; the processed image features and the image features before processing are connected by residuals, which provides a more direct path for the subsequent back propagation of the gradient, so that the gradient can be more effectively transmitted to each layer, so that the network can converge more quickly during training, reducing the number of iterations required for training, thereby improving the training efficiency. At this point, the final template image features and intermediate search area features are obtained.
[0066] In this step, the purpose of spatial channel summation attention is to quickly focus attention on the area of interest and obtain a larger receptive field, and spatial channel summation attention is achieved through convolution operations, which effectively avoids complex operations such as matrix multiplication and Softmax in pure ViT, greatly reduces the computational complexity, and achieves efficient self-attention.
[0067] Step 4: preprocess the template image of the first frame and the search area image of the subsequent frames, perform image block embedding operations on the template image and the search area image respectively, and encode the embedding position of each image block to obtain template features and search area features;
[0068] Input the template features and search area features into the two branches of the pre-trained dual-branch feature extraction network respectively, and repeat step 3 several times in an iterative manner to obtain the final template features and intermediate search features;
[0069] As a further preferred embodiment of the present invention, preprocessing the template image of the first frame and the search area image of the subsequent frames specifically includes the following steps:
[0070] Step 4.1, cutting the input image into a number of non-overlapping image blocks of fixed size;
[0071] Step 4.2, flatten each image block into a one-dimensional vector;
[0072] Step 4.3, project each flattened image block into a high-dimensional space through a linear transformation layer to convert it into an embedding vector of fixed dimension to obtain a preprocessed image;
[0073] Step 4.4, repeating steps 4.1 to 4.3 for the template image of the first frame and the search area image of the subsequent frames respectively, to obtain a preprocessed template image and a preprocessed search area image.
[0074] Step 5: Input the final template features and the intermediate search features into the feature fusion network for fusion to obtain a two-dimensional feature map, and input the two-dimensional feature map into the head network to obtain the target tracking frame;
[0075] See also Figure 3 As a further preferred embodiment of the present invention, the final template features and the intermediate search features are input into the feature fusion network for fusion, and obtaining the two-dimensional feature map specifically includes the following steps:
[0076] The final template features and the intermediate search features are concatenated to obtain the initial feature sequence which is passed into the feature fusion network;
[0077] The final template features and the intermediate search area features are fused using the Transformer-based variant self-attention to obtain the fusion result.
[0078] The search area part of the fusion result of the current layer feature fusion network is taken in series with the final template feature as the input of the next layer feature fusion network. After passing through several layers of feature fusion networks in an iterative manner, a two-dimensional feature map is obtained.
[0079] The template features and the intermediate search area features are fused using the Transformer-based variant self-attention, and the corresponding process of the fusion result is as follows:
[0080] ;
[0081] ;
[0082] ;
[0083] in, represents the matrix concatenation operation, Respectively represent the query of the search area feature, the key of the search area feature, the key of the template feature, the value of the search area feature, and the value of the template feature; ; Respectively The corresponding weight matrix is, Represent the second key and the second value respectively, ; Respectively The transpose of represents the normalized exponential activation function, which converts the similarity scores between the query and the key into a probability distribution, which is used to represent the attention weight of each element relative to other elements; represents the feature dimension, represents the scaling factor, divided by This is because the result of the dot product is usually very large, so the softmax result cannot express the similarity score well, and the gradient may disappear. In this case, we divide it by a scaling factor , which can alleviate this situation to a certain extent; represents the fusion result, Represent the final template features and the intermediate search area features respectively.
[0084] In this step, the query of variant self-attention comes only from the search area features, while the key and value are obtained by splicing the search area features and the template features. The purpose of variant self-attention is to extract and interact the obtained template features and search area features to obtain better feature fusion. The dimensional change of spliced keys and values has a significant impact on the attention calculation. It increases the computational complexity, but also enables the model to capture and fuse information from different angles, thereby improving the model's expressiveness and generalization performance. It allows the model to capture information from different subspaces in parallel, enhancing the model's ability to learn complex dependencies. This mechanism can help the model learn the complex relationships between different regions in the image.
[0085] Step 6: Using the training set as input, repeat steps 3 to 5 to train the target tracking model to obtain a trained target tracking model;
[0086] Step 7: Use the trained target tracking model to track the target.
[0087] The head network consists of three convolutional branches: the first branch performs center classification to generate a center score map, where each number represents the confidence that the center of the object is located at the corresponding predicted position, the second branch performs offset regression to calculate the discretization error of the center, and the third branch performs size regression to obtain the height and width of the object.
[0088] Since the head network consists of three convolution branches, the present invention designs a targeted training process, which specifically includes the following steps:
[0089] Get the training set, use the training set as input and repeat steps 3 to 5 to get the target tracking frame;
[0090] Based on the cross entropy loss, the first loss function is constructed using the target tracking box and the given true label. The corresponding process has the following relationship:
[0091] ;
[0092] in, represents the first loss function, which is used to determine the inconsistency between the prediction result and the true label. The image features are passed through the classification branch in the prediction head to obtain the centrality score map, where the position with the highest confidence is selected as the position of the predicted target; The true label of the target position is 1 or 0, where 1 or 0 indicates that the position is the target position or the background area respectively; represents the logarithmic function, Represents the probability value of the target location;
[0093] Based on the GIoU loss, the second loss function is constructed using the reference window of the target tracking box and the reference window of the given true value bounding box. The corresponding process has the following relationship:
[0094] ;
[0095] in, represents the second loss function, which is used to determine the positional relationship between the predicted bounding box and the true bounding box. The image features are processed through the size regression branch to obtain the width and height of the final bounding box. The reference window representing the ground truth bounding box, , Represents the center position coordinates of the true value window, Respectively represent the width and height of the true value window; represents the reference window for the predicted bounding box, , Represents the center coordinates of the prediction window, Respectively represent the width and height of the prediction window; express and The minimum circumscribed matrix of Represents the intersection-over-union calculation operation between the predicted box and the true box.
[0096] based on Loss, the third loss function is constructed by using the center position coordinates of the target tracking box reference window and the center position coordinates of the given true value bounding box reference window. The corresponding process has the following relationship:
[0097] ;
[0098] in, Represents the third loss function, which is used to calculate the mean absolute error between the predicted bounding box and the true bounding box center position. The image features are passed through the offset regression branch to obtain the final target center coordinates. The final predicted coordinates of the bounding box are calculated using the target center coordinates and height and width obtained by regression;
[0099] The total loss function is constructed using the first loss function, the second loss function, and the third loss function. The corresponding process has the following relationship:
[0100] ;
[0101] in, represents the total loss function, Respectively represent the weight hyperparameters corresponding to the second loss function and the third loss function, which are used to adjust the weights of different loss items in the total loss to reflect the importance of different tasks or balance the impact of different losses on model training;
[0102] The object tracking model is trained by updating the weights and learning parameters to minimize the total loss.
[0103] Please refer to Figure 4 This embodiment further provides a target tracking system based on spatial channel summation and attention, wherein the system applies the target tracking method based on spatial channel summation and attention as described above, and the system includes:
[0104] Building blocks for:
[0105] A spatial channel sum attention module is built based on the spatial attention mechanism and the channel attention mechanism. A dual-branch feature extraction network is built based on the spatial channel sum attention module and the feature integration module. A feature fusion network is built based on the variant self-attention mechanism. The dual-branch feature extraction network, the feature fusion network and the head network constitute the target tracking model.
[0106] Training modules for:
[0107] Using a large-scale data set to pre-train a dual-branch feature extraction network and a feature fusion network, a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network are obtained;
[0108] Extraction module 1, used for:
[0109] The image features are input into the pre-trained dual-branch feature extraction network, and the image features are fused using the feature integration module to obtain fused image features;
[0110] Perform residual connection on the image features and fused image features, and perform layer normalization. Then, send the layer normalization results to the spatial channel sum attention module for information interaction between the spatial domain and the channel domain to obtain the attention output.
[0111] The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map;
[0112] Extraction module 2, used for:
[0113] Preprocessing the template image of the first frame and the search area image of the subsequent frames, performing image block embedding operations on the template image and the search area image respectively, and encoding the embedding position of each image block to obtain template features and search area features;
[0114] The template features and search area features are respectively input into the two branches of the pre-trained dual-branch feature extraction network, and the operation of a part of the extraction module is repeated several times in an iterative manner to obtain the final template features and intermediate search features;
[0115] Compute module for:
[0116] The final template features and the intermediate search features are input into the feature fusion network for fusion to obtain a two-dimensional feature map, and the two-dimensional feature map is input into the head network to obtain the target tracking frame;
[0117] Learning modules for:
[0118] The target tracking model is trained by using the training set as input to obtain a trained target tracking model;
[0119] Tracking module for:
[0120] Use the trained target tracking model to perform target tracking.
[0121] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0122] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0123] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0124] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A target tracking method based on spatial channel summation attention, characterized in that: The method comprises the following steps: Step 1: Based on the spatial attention mechanism and the channel attention mechanism, a spatial channel sum attention module is constructed. Based on the spatial channel sum attention module and the feature integration module, a dual-branch feature extraction network is constructed. Based on the variant self-attention mechanism, a feature fusion network is constructed. The dual-branch feature extraction network, the feature fusion network and the head network constitute a target tracking model. Step 2: Pre-train the dual-branch feature extraction network and the feature fusion network using a large-scale data set to obtain a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network; Step 3: Input the image features into the pre-trained dual-branch feature extraction network, and use the feature integration module to fuse the image features to obtain fused image features; The image features and fused image features are residually connected and layer normalized. The layer normalization results are then sent to the spatial channel sum attention module for information interaction in the spatial domain and channel domain to obtain the attention output. The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map; Step 4: preprocess the template image of the first frame and the search area image of the subsequent frames, perform image block embedding operations on the template image and the search area image respectively, and encode the embedding position of each image block to obtain template features and search area features; Input the template features and search area features into the two branches of the pre-trained dual-branch feature extraction network respectively, and repeat step 3 several times in an iterative manner to obtain the final template features and intermediate search features; Step 5: Input the final template features and the intermediate search features into the feature fusion network for fusion to obtain a two-dimensional feature map, and input the two-dimensional feature map into the head network to obtain the target tracking frame; Step 6: Using the training set as input, repeat steps 3 to 5 to train the target tracking model to obtain a trained target tracking model; Step 7: Use the trained target tracking model to track the target.
2. The target tracking method based on spatial channel summation attention according to claim 1 is characterized in that: In step 1, the dual-branch feature extraction network includes a template branch feature extraction network for processing the template image and a search branch feature extraction network for processing the search area image. The template branch feature extraction network is set to 9 layers, and the search branch feature extraction network is set to 6 layers.
3. The target tracking method based on spatial channel summation attention according to claim 2 is characterized in that: In the step 1, the feature integration module is composed of a first 1×1 convolution, a normalization layer, a GeLU activated depthwise separable convolution layer, and a second 1×1 convolution, which are arranged in sequence.
4. The target tracking method based on spatial channel summation attention according to claim 3 is characterized in that: In step 3, the layer normalization result is sent to the spatial channel sum attention module for information interaction in the spatial domain and the channel domain, and obtaining the attention output specifically includes the following steps: Normalizing the layer results in generating a first query, a first key, and a first value; Using spatial attention for the first query and the first key, respectively, to assign different weights to different spatial positions of the input feature map, and obtain the spatial attention outputs of the first query and the first key; The spatial attention outputs of the first query and the first key are respectively assigned different weights to different channels of the input feature map using channel attention to obtain the channel attention outputs of the first query and the first key; Then add the channel attention outputs of the first query and the first key, perform a linear transformation to integrate the context information, and then perform a dot product operation on the linear transformation result and the first value to obtain the attention output.
5. The target tracking method based on spatial channel summation attention according to claim 4 is characterized in that: In step 3, the layer normalization result is sent to the spatial channel sum attention module for information interaction between the spatial domain and the channel domain. The process of obtaining the attention output corresponds to the following relationship: ; ; in, Represent the first query, first key, and first value respectively. , , , Indicates the number of image blocks, Represents the feature dimension of each image block; Respectively represent the weight matrices corresponding to the first query, the first key, and the first value; represents the linear transformation operation used to integrate context information, represents the context mapping function, represents the spatial attention calculation operation, represents the channel attention calculation operation, represents the attention output, Embedded representation of the input sequence of the spatial channel sum attention module, i.e., layer normalization result.
6. The target tracking method based on spatial channel summation attention according to claim 5 is characterized in that: In step 4, preprocessing the template image of the first frame and the search area image of the subsequent frames specifically includes the following steps: Step 4.1, cutting the input image into a number of non-overlapping image blocks of fixed size; Step 4.2, flatten each image block into a one-dimensional vector; Step 4.3, project each flattened image block into a high-dimensional space through a linear transformation layer to convert it into an embedding vector of fixed dimension to obtain a preprocessed image; Step 4.4, repeating steps 4.1 to 4.3 for the template image of the first frame and the search area image of the subsequent frames respectively, to obtain a preprocessed template image and a preprocessed search area image.
7. The target tracking method based on spatial channel summation attention according to claim 6 is characterized in that: In step 5, the final template features and the intermediate search features are input into the feature fusion network for fusion, and obtaining the two-dimensional feature map specifically includes the following steps: The final template features and the intermediate search features are concatenated to obtain the initial feature sequence which is passed into the feature fusion network; The final template features and the intermediate search area features are fused using the Transformer-based variant self-attention to obtain the fusion result. The corresponding process has the following relationship: ; ; ; in, represents the matrix concatenation operation, Respectively represent the query of the search area feature, the key of the search area feature, the key of the template feature, the value of the search area feature, and the value of the template feature; ; Respectively The corresponding weight matrix is, Represent the second key and the second value respectively, ; Respectively The transpose of represents the normalized exponential activation function, represents the feature dimension, represents the scaling factor, represents the fusion result, Represent the final template features and the intermediate search area features respectively; The search area part of the fusion result of the current layer feature fusion network is taken in series with the final template feature as the input of the next layer feature fusion network. After passing through several layers of feature fusion networks in an iterative manner, a two-dimensional feature map is obtained.
8. The target tracking method based on spatial channel summation attention according to claim 7 is characterized in that: The head network consists of three convolutional branches: the first branch performs center classification to generate a center score map, where each number represents the confidence that the center of the object is located at the corresponding predicted position, the second branch performs offset regression to calculate the discretization error of the center, and the third branch performs size regression to obtain the height and width of the object.
9. The target tracking method based on spatial channel summation attention according to claim 8, characterized in that: In step 6, the target tracking model is trained by repeating steps 3 to 5 using the training set as input, and obtaining the trained target tracking model specifically includes the following steps: Get the training set, use the training set as input and repeat steps 3 to 5 to get the target tracking frame; Based on the cross entropy loss, the first loss function is constructed using the target tracking box and the given true label. The corresponding process has the following relationship: ; in, represents the first loss function, represents the true label of the target location, represents the logarithmic function, Represents the probability value of the target location; Based on the GIoU loss, the second loss function is constructed using the reference window of the target tracking box and the reference window of the given true value bounding box. The corresponding process has the following relationship: ; in, represents the second loss function, The reference window representing the ground truth bounding box, , Represents the center position coordinates of the true value window, Respectively represent the width and height of the true value window; represents the reference window for the predicted bounding box, , Represents the center coordinates of the prediction window, Respectively represent the width and height of the prediction window; express and The minimum circumscribed matrix of Represents the intersection-and-union calculation operation between the predicted box and the true box; based on Loss, the third loss function is constructed by using the center position coordinates of the target tracking box reference window and the center position coordinates of the given true value bounding box reference window. The corresponding process has the following relationship: ; in, represents the third loss function; The total loss function is constructed using the first loss function, the second loss function, and the third loss function. The corresponding process has the following relationship: ; in, represents the total loss function, Respectively represent the weight hyperparameters corresponding to the second loss function and the third loss function; The object tracking model is trained by updating the weights and learning parameters to minimize the total loss.
10. A target tracking system based on spatial channel summation attention, characterized in that: The system applies the target tracking method based on spatial channel summation attention according to any one of claims 1 to 9, and the system comprises: Building blocks for: A spatial channel sum attention module is built based on the spatial attention mechanism and the channel attention mechanism. A dual-branch feature extraction network is built based on the spatial channel sum attention module and the feature integration module. A feature fusion network is built based on the variant self-attention mechanism. The dual-branch feature extraction network, the feature fusion network and the head network constitute the target tracking model. Training modules for: Using a large-scale data set to pre-train a dual-branch feature extraction network and a feature fusion network, a pre-trained dual-branch feature extraction network and a pre-trained feature fusion network are obtained; Extraction module 1, used for: The image features are input into the pre-trained dual-branch feature extraction network, and the image features are fused using the feature integration module to obtain fused image features; The image features and fused image features are residually connected and layer normalized. The layer normalization results are then sent to the spatial channel sum attention module for information interaction in the spatial domain and channel domain to obtain the attention output. The attention output is passed through layer normalization and multi-layer perceptron in sequence, and then the multi-layer perceptron output is residually connected with the attention output to obtain a feature map; Extraction module 2, used for: Preprocessing the template image of the first frame and the search area image of the subsequent frames, performing image block embedding operations on the template image and the search area image respectively, and encoding the embedding position of each image block to obtain template features and search area features; The template features and search area features are respectively input into the two branches of the pre-trained dual-branch feature extraction network, and the operation of a part of the extraction module is repeated several times in an iterative manner to obtain the final template features and intermediate search features; Compute module for: The final template features and the intermediate search features are input into the feature fusion network for fusion to obtain a two-dimensional feature map, and the two-dimensional feature map is input into the head network to obtain the target tracking frame; Learning modules for: The target tracking model is trained by using the training set as input to obtain a trained target tracking model; Tracking module for: Use the trained target tracking model to perform target tracking.
Citation Information
Patent Citations
Target tracking method and system based on dual attention feature fusion network
CN116030097A
False face detection method based on frequency attention feature fusion, medium and equipment
CN116434351A