Visual target tracking system and method based on space-time attention
By using the combination of space-time excitation module, motion excitation module and channel excitation module in the visual target tracking system, combined with the sparse Transformer network, the problem of target tracking offset in the prior art under complex background is solved, and a high-precision and robust visual target tracking effect is achieved.
Patent Information
- Application Number
- CN202411751268.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-02
AI Technical Summary
Existing visual target tracking algorithms are difficult to effectively utilize the spatial and temporal dynamic correlation of videos in complex contexts, resulting in target tracking being easily offset.
The visual target tracking system based on space-time attention is adopted, and the space-time information in the video sequence is modeled through the space-time excitation module, the motion excitation module and the channel excitation module, and the three complementary features are fused. The feature information is integrated and the spatiotemporal feature fusion network enhanced by the sparse Transformer is input to learn discriminant space-time, motion and spatial information.
It effectively improves the performance of the tracker, and can maintain high accuracy and robustness in complex environments such as lighting changes, scale changes, occlusion, background blur and deformation, achieving real-time tracking effects.
Smart Images

Figure CN119992123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, in particular to the field of visual tracking technology of digital images, and in particular to a visual target tracking system and method based on spatiotemporal attention. Background Art
[0002] Visual object tracking is an important research topic in the field of computer vision, and has a wide range of applications in video understanding, human-computer interaction, visual surveillance, and autonomous driving. In recent years, object tracking based on twin networks has become increasingly popular. The SiamFC algorithm combines CNN features with the Siamese framework to implement a fully convolutional twin network tracking algorithm. Since then, a series of improved algorithms based on Siamese network trackers have emerged. Compared with SiamFC's multi-scale estimation, SiamRPN and SiamFC++ use anchor or anchor-free bounding box estimation mechanisms to effectively improve the positioning accuracy of the target; at the same time, SiamRPN++ and Ocean use the powerful ResNet50 as the backbone to enhance feature expression capabilities; in addition, algorithms such as ATOM and DiMP combine online updates with Siamese networks to enable trackers to achieve better tracking performance.
[0003] CNN-based visual features have been successful in a variety of image tasks, but due to the diversity and complexity of video content, extracting robust spatiotemporal features from video sequences is still a challenge. Spatiotemporal feature extraction networks are widely used in behavior recognition and action understanding. In tracking tasks, since videos have strong spatiotemporal dynamic correlations, spatiotemporal feature extraction networks can be considered. However, the feature extraction methods of existing tracking algorithms mostly rely on image-level features in the video, performing two-dimensional convolution on an independent frame in the video. This method does not effectively utilize the spatiotemporal dynamic correlations of the video. Figure 1 As shown in Image (a), the target foreground and background are less distinguishable; Figure 1 As shown in the image (b), the target in the bounding box occupies 20% to 30% of the box, and interference and background clutter appear in the response graphs of the DiMP and TrDiMP algorithms. This shows that in this complex background, it is easy to offset the tracking process by only using the surface features of the target in an independent frame. Therefore, it is necessary to use the network to learn the spatiotemporal dynamic changes of the target in the video. In response to this type of problem, this paper characterizes the temporal and spatial information of the target in the video by dividing the video feature information into three parts: spatial information flow, temporal information, and motion information flow.
[0004] Transformer networks are widely used in various video tasks and have shown good learning and generalization capabilities in different scenarios. Algorithms such as TrDiMP and TrTr have successfully applied the Transformer structure in visual object tracking and achieved good results. However, due to the limitation of computational complexity, Transformer is not good at processing longer contexts. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a visual target tracking method based on spatiotemporal attention, which can learn discriminative spatiotemporal, motion and spatial information by exploring the spatiotemporal contextual relationship of consecutive frames, thereby effectively improving the performance of the tracker.
[0006] The technical solution of the present invention is: the present invention is a visual target tracking system based on spatiotemporal attention, which is special in that: the visual target tracking system based on spatiotemporal attention includes an image acquisition module, a feature extraction module, a spatiotemporal feature enhancement module and a classification regression positioning module connected in sequence, the image acquisition module is used to acquire multiple frames of images in a video sequence as training images and test images respectively; the feature extraction module is used to input the training images and test images into a visual tracking network for feature extraction to obtain feature representation of the video sequence; the spatiotemporal feature enhancement module includes a spatiotemporal excitation module, a motion excitation module, a channel excitation module and a sparse Transformer, the sparse Transformer module includes a sparse Transformer encoder and a sparse Transformer decoder connected to each other, and the feature extraction module is respectively connected through the spatiotemporal excitation module, the motion excitation module and The channel excitation module is connected to the sparse Transformer encoder, and the sparse Transformer encoder and the sparse Transformer decoder are respectively connected to the classification regression and positioning module. The spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module. Then the three mixed complementary features are fused, and the integrated feature information is input into the spatiotemporal feature fusion network enhanced by the sparse Transformer module to model the spatiotemporal information; the classification regression and positioning module includes a model predictor f and conv calculation. The model predictor f iteratively calculates and optimizes the aggregated features of the sparse Transformer decoder, and then performs conv calculation on the output weights of the decoder output. These weights are used for the target classification operation of the feature map; finally, the bounding box is obtained by the predicted offset in the classification map and the coordinates of the best prediction point.
[0007] Furthermore, the spatiotemporal excitation module uses single-channel 3D convolution to represent spatiotemporal information. Compared with traditional 3D convolution operations, STE is more computationally efficient and can perceive spatiotemporal information from a fine feature excitation from each channel of the input feature.
[0008] The motion excitation module calculates the time difference between adjacent frames, and then uses these time differences to excite the motion sensitive channel; it mainly describes the movement of the action between each two adjacent frames, subtracts the latter frame from the previous frame, and splices them in dimension, mainly based on the frame difference method to calculate the motion characteristics of adjacent frames.
[0009] The channel excitation module adaptively recalibrates the channel feature response by explicitly modeling the interdependence between channels in time. The video sequence of the channel excitation module contains temporal information, and a 1D convolution in the time domain is inserted between the squeeze and unsqueeze of the feature channel to enhance the interdependence of the channel in the time domain, and the excitation space information of the input feature is multiplied by the attention matrix.
[0010] The present invention also provides a method for implementing a visual target tracking system based on spatiotemporal attention, the special feature of which is that the method comprises the following steps:
[0011] 1) Acquire multiple frames of images in the video sequence as training images and test images respectively through image acquisition;
[0012] 2) Input the training image and the test image into the visual tracking network through the feature extraction module to extract features and obtain the feature representation of the video sequence;
[0013] 3) The spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module. Then the three hybrid complementary features are fused and the integrated feature information is input into the sparse Transformer enhanced spatiotemporal feature fusion network to model the spatiotemporal information.
[0014] 4) The target classification operation of the integrated feature map is performed through the classification regression positioning module, and finally the bounding box is obtained through the predicted offset in the classification map and the coordinates of the best prediction point.
[0015] Furthermore, in step 3), the spatiotemporal excitation module mainly uses a single-channel 3D convolution to represent the spatiotemporal information, splits the interaction between time and space, and only calculates a separate Attention for the interaction in space, and then calculates the Attention again in time. This makes it possible to obtain a spatiotemporal MaskM (N×T×1×H×W) with a very small amount of calculation, and multiply the input features with the attention map to obtain the corresponding features excited by the spatiotemporal information.
[0016] Furthermore, the motion excitation module in step 3) describes the movement of the action between each two adjacent frames, where the input is a 5D tensor (N, T, C, H, W); N corresponds to the batch, T is the number of video frames, C is the number of channels, and H and W correspond to the feature map height and width respectively; first, the input is subjected to a 1×1 convolution to perform a channel dimension reduction operation, and the five-dimensional tensor is separated into T four-dimensional tensors in the T dimension, where the latter T-1 tensors pass through a shared 2D convolution kernel, and then the latter frame is subtracted from the previous frame and concatenated in the T dimension, and the motion features of each frame are calculated according to formula (1), where the modeling of the motion features is expressed as:
[0017] F m =K*F r [:,t+1,:,:,:]-F r [:,t+1,:,:,:] (1)
[0018] Where K is a 3x3 convolution. First, the input is squeezed through a 1×1 convolution. The motion features of each frame are calculated according to the above formula, and 0 is filled into the last position, which is expressed as: F M =[F m (1),…,F m (t-1),0], where F M The size is T×c / r×H×W, then F M After spatial average pooling and 1×1 convolution dimensionality reduction operations, the mask M is obtained after Sigmod activation.
[0019] Furthermore, in step 3), the video sequence of the channel excitation module contains temporal information, and a 1D convolution in the time domain is inserted between the squeeze and unsqueeze of the channel to enhance the interdependence of the channels in the time domain; the channel excitation module calculates a channel-based attention matrix, and excites the spatial information of the input features by multiplying the attention matrix; in addition, the channel excitation module inserts a 1×1 convolution layer between the two FC layers to describe the temporal information of the channel features;
[0020] Furthermore, the network structure of the sparse Transformer in step 3) is composed of a sparse attention module, a multi-layer perceptron (MLP), and a position convolutional coding layer. The self-attention function operates on the query Q, key K, and value V in a proportional point generation manner. The process can be expressed as:
[0021]
[0022] Where dx is the key dimension of the attention matrix, and SAttention is sparse self-attention. The feature is added to the Q key value before the backbone network extracts it and inputs it into the Transformer. Then, the representation ability of the model is enhanced by expanding the attention mechanism to multiple heads. The formula is as follows:
[0023] MultiHead(Q i ,K i ,V i )=Concat(H1,…,H n )W O (4)
[0024] H i = SAttention(QW i Q ,KW i K ,VW i V ) (5)
[0025] In the above formula, W i Q ,W i K ,W i V ∈ is the parameter of linear projection, n is the number of attention heads, and after Add&Norm, it is LayerNorm(x+Linerlayer(x)), where Norm is the smooth input feature;
[0026] Position encoding is used in self-attention; then the calculated x is input into the MLP layer and the position convolutional encoding layer, and its formula is shown as follows;
[0027] PE(Q,K,V)=Norm(x+MLP(x)) (6)
[0028]
[0029] Among them, MLP is a fully connected feedforward network that transforms the original features into arbitrary dimensions. Formula (7) is a position convolutional coding layer that encodes the image to retain the detailed information of the local features of the target.
[0030] The present invention provides a visual target tracking system based on spatiotemporal attention. The three modules of the spatial-temporal excitation module (Spatial-Temporal Excitation, STE), the motion excitation module (Motion Excitation, ME) and the channel excitation module (Channel Excitation, CE) are designed to model the temporal and spatial information in the video sequence and then fuse the three complementary features to effectively aggregate the temporal and spatial information. The spatial-temporal excitation module uses a single-channel 3D convolution to characterize the spatiotemporal information. The motion excitation module calculates the time difference between adjacent frames and then uses these time differences to excite motion-sensitive channels. The channel excitation module adaptively recalibrates the channel feature response by explicitly modeling the interdependence between channels in time, making the greatest use of the dependency between channels.
[0031] The present invention also relates to a sparse Transformer structure, which processes the data encoded by the three streams of time, space, motion and space, and adopts a sparse coding method to reduce the computational complexity and provide fine features in the convolution position coding layer: CNN effectively retains local features through convolution but cannot capture global dependencies. The present invention uses multi-frame input to extract time sequence information using a spatiotemporal network structure, and then uses the self-attention of the sparse Transformer to capture long-distance dependencies to increase the stability of the output results, better handle occluded scenes, and predict the movement of the target. However, the linear calculation of the Transformer will ignore the details of local features, which will reduce the discernibility between the target background and foreground. Therefore, relative to the classic Transformer structure, the present invention proposes a network structure that can combine the local features of CNN with the global features of the sparse Transformer to model mixed spatiotemporal information. The present invention extracts spatiotemporal features as the input of the Transformer coding structure, and proposes a coding layer and a decoding layer of a sparse Transformer. The encoding layer of the Transformer is used to encode the image. In addition, linear operations during the encoding process will damage the position information within each patch, making it more difficult to extract local information within the patch. The present invention designs a multi-layer perceptron (MLP) and a position convolutional coding layer in the Transformer to flexibly handle the position coding layer with varying resolutions, and then enhances the spatiotemporal feature representation capabilities through the Transformer encoder and decoder. The sparse Transformer structure captures long-range dependencies through the Transformer attention mechanism; finally, in the sparse Transformer encoder and decoder, the position convolutional coding layer is used to provide accurate position encoding for feature fusion, and efficient spatiotemporal information is learned by constructing a Transformer-enhanced spatiotemporal feature fusion network.
[0032] The present invention also proposes a visual target tracking algorithm based on spatiotemporal attention. The spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module, and then the three hybrid complementary features are fused to effectively aggregate the time, space and motion information. Efficient spatiotemporal information is learned by constructing a spatiotemporal feature fusion network enhanced by Transformer. It is proved that the spatiotemporal information can be effectively modeled by constructing a spatiotemporal network enhanced by context information, so that the tracker can obtain good tracking performance.
[0033] Therefore, compared with the prior art, the present invention has the following advantages:
[0034] 1) The present invention can achieve real-time tracking effect, that is, it has certain economic benefits.
[0035] 2) In complex and changeable tracking environments such as illumination changes, scale changes, occlusions, background blur and deformation, the present invention can still maintain high accuracy and robustness and has obvious advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is the response diagram result analysis diagram in the background technology;
[0037] Figure 2 is a system block diagram of the present invention;
[0038] Figure 3 In the method of the present invention, the network structure diagram of the spatiotemporal excitation module, the motion excitation module and the channel excitation module;
[0039] Figure 4 It is a sparse Transformer network structure diagram in the method of the present invention;
[0040] Figure 5 It is the experimental result figure of the method of the present invention;
[0041] Figure 6 It is a heat map visualization diagram of the method of the present invention;
[0042] Figure 7 It is a radar chart of the method of the present invention under different attributes;
[0043] Figure 8 This is a comparison diagram of the results of the VOT experiment using the method of the present invention. DETAILED DESCRIPTION
[0044] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0045] See also Figure 2The structure of the specific embodiment of the visual target tracking system based on spatiotemporal attention provided by the present invention mainly includes four parts: image acquisition module, feature extraction module, spatiotemporal feature enhancement module and classification regression positioning module. The image acquisition module is used to acquire multiple frames of images in the video sequence as training images and test images respectively; the feature extraction module is used to input the training images and test images into the visual tracking network for feature extraction to obtain the feature representation of the video sequence; the spatiotemporal feature enhancement module includes a spatiotemporal excitation module, a motion excitation module and a channel excitation module. The spatiotemporal information in the video sequence is modeled by the spatiotemporal excitation module, the motion excitation module and the channel excitation module at the same time, and then the three mixed complementary features are fused to effectively aggregate the time and space information; the spatiotemporal feature fusion network enhanced by the sparse Transformer is used to model the spatiotemporal information by integrating the feature information; the classification regression positioning module is used to perform the target classification operation on the integrated feature map. Finally, the bounding box is obtained by the predicted offset in the classification map and the coordinates of the best prediction point.
[0046] Among them, the spatiotemporal excitation module uses single-channel 3D convolution to represent spatiotemporal information; compared with traditional 3D convolution operations, STE has higher computational efficiency and can perceive spatiotemporal information from a fine feature excitation from each channel of the input feature.
[0047] The motion excitation module mainly describes the movement of actions between each two adjacent frames. It subtracts the latter frame from the previous frame and splices them in dimension. It mainly calculates the motion features of adjacent frames based on the frame difference method; calculates the time difference between adjacent frames, and then uses these time differences to excite the motion sensitive channel.
[0048] The channel excitation module adaptively recalibrates the channel feature response by explicitly modeling the interdependence between channels in time, making the most of the inter-channel dependencies. The video sequence of the channel excitation module contains temporal information, and a 1D convolution in the time domain is inserted between the squeeze and unsqueeze of the feature channel to enhance the interdependence of the channels in the time domain, and the excitation space information of the input features is multiplied by the attention matrix.
[0049] When the target is in a complex environment, in order to capture the target's spatiotemporal, motion and spatial information, the spatiotemporal excitation module, motion excitation module and channel excitation module are used to construct a hybrid feature with three complementary information. The fused feature map encodes the target's time, motion and appearance features. These features are used as input to enhance the fused features through the Transformer encoder and decoder.
[0050] The classification regression positioning module includes a model predictor f and conv calculation. First, the model predictor f is initialized by the target area features. The initialized model predictor f is combined with the sparse Transformer encoder aggregation features to be optimized in an iterative manner. Then, f is conv calculated with the sparse Transformer decoder to output weights. These weights are used for the target classification and positioning operation of the feature map.
[0051] The present invention uses a multi-frame video sequence as input and uses the ResNet-50 backbone network to extract deep features. The feature map is input into the Transformer encoder and decoder, and the encoding result is input into the model predictor f composed of the optimizer. The model predictor f outputs the weights of the convolutional layer, which are used in the target classification operation of the feature map extracted from the test frame. Finally, the bounding box is obtained by the predicted offset in the classification map and the coordinates of the best prediction point.
[0052] The present invention provides a method for visual target tracking based on spatiotemporal attention, comprising the following steps:
[0053] 1) Acquire multiple frames of images in the video sequence as training images and test images respectively through image acquisition;
[0054] 2) Input the training image and the test image into the visual tracking network through the feature extraction module to extract features and obtain the feature representation of the video sequence;
[0055] 3) The spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module. Then the three hybrid complementary features are fused and the integrated feature information is input into the sparse Transformer enhanced spatiotemporal feature fusion network to model the spatiotemporal information.
[0056] 4) The target classification operation of the integrated feature map is performed through the classification regression positioning module, and finally the bounding box is obtained through the predicted offset in the classification map and the coordinates of the best prediction point.
[0057] See also Figure 3 The spatiotemporal excitation module mainly uses a single-channel 3D convolution to represent spatiotemporal information. This structure splits the interaction between time and space, interacts in space, only calculates a separate Attention, and then calculates Attention again in time, which makes it possible to obtain a spatiotemporal MaskM(N×T×1×H×W) with a very small amount of calculation. The input features are dot-multiplied with the attention map to obtain the corresponding features stimulated by the spatiotemporal information. Compared with traditional 3D convolution operations, STE is more computationally efficient because the features F input to the 3D convolution are averaged across channels. Each channel of the input X can perceive the importance of spatiotemporal information from a fine feature excitation M.
[0058] The motion excitation module ME module mainly describes the movement of the action between each two adjacent frames, where the input is a 5D tensor (N, T, C, H, W). N corresponds to the batch, T is the number of video frames, C is the number of channels, and H and W correspond to the feature map height and width respectively. First, the input is subjected to a 1×1 convolution to perform a channel dimension reduction operation, and the five-dimensional tensor is separated into T four-dimensional tensors in the T dimension. The latter T-1 tensors are passed through a shared 2D convolution kernel, and then the latter frame is subtracted from the previous frame and spliced in the T dimension. The motion features of each frame are calculated according to formula (1), where the modeling of the motion features is expressed as follows:
[0059] F m =K*F r [:,t+1,:,:,:]-F r [:,t+1,:,:,:] (1)
[0060] Where K is a 3x3 convolution. First, the input is squeezed after a 1×1 convolution. The motion features of each frame are calculated according to the above formula, and 0 is filled to the last position, which is expressed as F. M =[F m (1),…,F m (t-1),0] where F M The size is T×c / r×H×W, then F M After spatial average pooling and 1×1 convolution dimensionality reduction operations, the mask M is obtained after Sigmod activation.
[0061] Channel Excitation Module Network Structure Video sequences contain temporal information, so a 1D convolution in the time domain is inserted between the squeeze and unsqueeze of the channel to enhance the interdependence of the channels in the time domain. By calculating a channel-based attention matrix. The excitation space information of the input features is multiplied by the attention matrix. In addition, the CE module inserts a 1×1 convolution layer between the two FC layers to describe the temporal information of the channel features.
[0062] See also Figure 4 , the sparse Transformer network structure of the present invention includes the following contents:
[0063] The sparse Transformer network structure is as follows Figure 4 As shown. This part is composed of a sparse attention module, a multi-layer perceptron (MLP), and a position convolutional coding layer. The self-attention function operates on the query Q, key K, and value V in a proportional point generation manner. The process can be expressed as:
[0064]
[0065] Where dx is the key dimension of the attention matrix, and SAttention is sparse self-attention. The feature is extracted by the backbone network and added to the Q key value before it is input into the Transformer. Then, by expanding the attention mechanism to multiple heads, the representation ability of the model is enhanced, and its formula is expressed as follows:
[0066] MultiHead(Q i ,K i ,V i )=Concat(H1,…,H n )W O (4)
[0067] H i = SAttention(QW i Q ,KW i K ,VW i V ) (5)
[0068] In the above formula, W i Q ,W i K ,W i V ∈ is the parameter of linear projection, n is the number of attention heads, and after Add&Norm, it is LayerNorm(x+Linerlayer(x)), where Norm is the smooth input feature;
[0069] The Patch Flatten operation will damage the position information within the patch, so we use position encoding in self-attention. Then the calculated x is input into the MLP layer and the position convolutional encoding layer, and its formula is shown as follows;
[0070] PE(Q,K,V)=Norm(x+MLP(x)) (6)
[0071]
[0072] Among them, MLP is a fully connected feedforward network that converts the original features into any dimension, and Equation 7 is a position convolutional coding layer that encodes the image to retain the target local feature detail information.
[0073] To verify the effectiveness of the method of the present invention, the algorithm of this chapter was implemented on an Inteli5-8400 CPU @ 2.80 GHz processor using Python programming under the Ubuntu operating system, and accelerated using a GPU (NVIDIA GTX 1080Ti).
[0074] The training data comes from the split of TrackingNet, COCO, LaSOT and GOT-10k datasets. The details of the experiment are as follows: ResNet-50 pre-trained on the ImageNet dataset is selected to initialize the model and the neural network. All experiments are trained for 50 epochs, using the ADAM optimizer to accelerate network convergence, and the initial learning rate of the backbone network is 10 -5 , the learning rate decreased by 10 over 10 epochs -1 .
[0075] See also Figure 5 ,Due to the complexity of underwater targets, the Turtle video sequence was selected to ,observe the experimental results.,In the video sequence, the foreground and background of the target are ,lowly distinguishable, and the target appears many times during the ,movement process surrounded by a large number of similar interferences.,Compared with other algorithms, this paper can effectively track the ,target. Figure 5 The second figure is a Skiing video sequence. Due to occlusion and target disappearance, the present invention can track the target better than the background technology, and when the target disappears and reappears, the method of the present invention can still track the target.
[0076] See also Figure 6 To observe the effect, the present invention visualizes the activation map of the training model backbone network and compares the heat map visualization of the background technology, the present invention including the convolution position encoding layer and the proposed algorithm without the encoding layer, such as Figure 6 As shown. Tr is the heat map visualization result of the standard Transformer, and PE is the heat map visualization result after introducing the convolutional position encoding layer. It can be seen from the figure that when the target appearance undergoes drastic deformation, the background technology has a higher activation value for the interference around the target, indicating that these channels pay more attention to the interference rather than the target itself, while the method of the present invention has a higher activation value in the target area, the temporal information pays more attention to the moving part, and the spatial information pays more attention to the fine spatial features of the complete part. This verifies that the backbone network effectively extracts complementary information of the target.
[0077] In order to verify the actual effect of the algorithm, LaSOT, VOT2019, and VOT2018-LT datasets are selected for evaluation. DiMP, PrDiMP, TrDiMP, TrTr, STMTrack and other background technologies are selected as reference methods.
[0078] See also Figure 7, the AUC (Area Under Curve) under 14 different attributes was tested on the LaSOT dataset and a radar chart was drawn. These attributes are: Camera Motion (CM), Target Deformation during Motion (View oint Change, VC), Fast Motion (FM) and other 14 attributes. In sports videos, low image resolution and target deformation often occur. In this type of video, the spatial features that can be extracted from a single frame image are very limited. The method of the present invention fully and effectively aggregates spatial information in the time dimension. Compared with the background technology, the AUC in motion blurred videos is improved by 3.2%; the AUC in fast motion videos is improved by 2.4%. Compared with other background technologies, it has obvious advantages in sports video sequences. It shows that the algorithm in this chapter has good robustness when dealing with complex scenes.
[0079] To further evaluate the effectiveness of the algorithm in visual tracking, the algorithm is evaluated on the VOT2019 dataset and compared with other background technologies. VOT mainly includes three indicators, EAO (Expected Average Overlap), Accuracy and Robustness, to analyze the tracking effect. The results are shown in Table 1. The best result is bold, the second best result is underlined, and the third best result is dotted underlined. Figure 8 , the right side shows the results of the top ten algorithms. Compared with the background technology, the method of the present invention improves EAO by 1.3% in the VOT2019 dataset, which further illustrates the effectiveness of the present invention.
[0080] Table 1 Comparison of VOT2019 algorithm experimental results
[0081]
[0082] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
[0083] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.
[0084] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.
Claims
1. A visual target tracking system based on spatiotemporal attention, characterized in that: The visual target tracking system based on spatiotemporal attention includes an image acquisition module, a feature extraction module, a spatiotemporal feature enhancement module and a classification regression positioning module connected in sequence, wherein the image acquisition module is used to acquire multiple frames of images in a video sequence as training images and test images respectively; the feature extraction module is used to input the training images and test images into a visual tracking network for feature extraction to obtain feature representations of the video sequence; the spatiotemporal feature enhancement module includes a spatiotemporal excitation module, a motion excitation module, a channel excitation module and a sparse Transformer, the sparse Transformer module includes a sparse Transformer encoder and a sparse Transformer decoder connected to each other, and the feature extraction module communicates with the sparse Transformer through the spatiotemporal excitation module, the motion excitation module and the channel excitation module respectively. r encoder, the sparse Transformer encoder and the sparse Transformer decoder are respectively connected to the classification regression positioning module, and the spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module, and then the three mixed complementary features are fused, and the integrated feature information is input into the spatiotemporal feature fusion network enhanced by the sparse Transformer module to model the spatiotemporal information; the classification regression positioning module includes a model predictor f and conv calculation, the model predictor f iteratively calculates and optimizes the sparse Transformer decoder aggregated features, and then conv calculates the output weights with the decoder output, and these weights are used for the target classification operation of the feature map; finally, the bounding box is obtained through the predicted offset in the classification map and the coordinates of the best prediction point.
2. The visual target tracking system based on spatiotemporal attention according to claim 1, characterized in that: The spatiotemporal excitation module uses single-channel 3D convolution to characterize spatiotemporal information; the motion excitation module calculates the time differences between adjacent frames and then uses these time differences to excite motion-sensitive channels; the channel excitation module adaptively recalibrates channel feature responses by explicitly modeling the interdependence between channels in time.
3. A method for implementing a visual target tracking system based on spatiotemporal attention, characterized in that: The method comprises the following steps: 1) Acquire multiple frames of images in the video sequence as training images and test images respectively through image acquisition; 2) Input the training image and the test image into the visual tracking network through the feature extraction module to extract features and obtain the feature representation of the video sequence; 3) The spatiotemporal information in the video sequence is modeled simultaneously through the spatiotemporal excitation module, the motion excitation module and the channel excitation module. Then the three hybrid complementary features are fused and the integrated feature information is input into the sparse Transformer enhanced spatiotemporal feature fusion network to model the spatiotemporal information. 4) The target classification operation of the integrated feature map is performed through the classification regression positioning module, and finally the bounding box is obtained through the predicted offset in the classification map and the coordinates of the best prediction point.
4. The method for visual target tracking based on spatiotemporal attention according to claim 3, characterized in that: In the step 3), the spatiotemporal excitation module mainly uses a single-channel 3D convolution to characterize the spatiotemporal information, splits the interaction between time and space, and only calculates a separate Attention for the interaction in space, and then calculates the Attention again in time. This makes it possible to obtain a spatiotemporal MaskM (N×T×1×H×W) with a very small amount of calculation, and obtain the corresponding features stimulated by the spatiotemporal information by multiplying the input features with the attention map.
5. The method for visual target tracking based on spatiotemporal attention according to claim 4, characterized in that: In step 3), the motion excitation module describes the movement of the action between each two adjacent frames, where the input is a 5D tensor (N, T, C, H, W); N corresponds to the batch, T is the number of video frames, C is the number of channels, and H and W correspond to the feature map height and width respectively; first, the input is subjected to a 1×1 convolution to perform a channel dimension reduction operation, and the five-dimensional tensor is separated into T four-dimensional tensors in the T dimension, where the latter T-1 tensors pass through a shared 2D convolution kernel, and then the latter frame is subtracted from the previous frame and spliced in the T dimension, and the motion features of each frame are calculated according to formula (1), where the modeling of the motion features is expressed as: F m =K*F r [:,t+1,:,:,:]-F r [:,t+1,:,:,:] (1) Where K is a 3x3 convolution. First, the input is squeezed through a 1×1 convolution. The motion features of each frame are calculated according to the above formula, and 0 is filled into the last position, which is expressed as: F M =[F m (1),…,F m (t-1),0], where F M The size is T×cr×H×W, then F M After spatial average pooling and 1×1 convolution dimensionality reduction operations, the mask M is obtained after Sigmod activation.
6. The method for visual target tracking based on spatiotemporal attention according to claim 5, characterized in that: The video sequence of the channel excitation module in step 3) contains timing information, and a 1D convolution in the time domain is inserted between the squeeze and unsqueeze of the channel to enhance the mutual dependence of the channels in the time domain; The channel excitation module calculates a channel-based attention matrix and multiplies the input feature excitation space information by the attention matrix; In addition, the channel excitation module inserts a 1×1 convolutional layer between two FC layers to describe the temporal information of channel features.
7. The method for visual target tracking based on spatiotemporal attention according to claim 6, characterized in that: The network structure of the sparse Transformer in step 3) is composed of a sparse attention module, a multi-layer perceptron (MLP) and a position convolutional coding layer. The self-attention function operates on the query Q, key K, and value V in a proportional point generation manner. The process can be expressed as: Where dx is the key dimension of the attention matrix, and SAttention is sparse self-attention. The feature is added to the Q key value before the backbone network extracts it and inputs it into the Transformer. Then, the representation ability of the model is enhanced by expanding the attention mechanism to multiple heads. The formula is as follows: MultiHead(Q i ,K i ,V i )=Concat(H1,…,H n )W O (4) H i =SAttention(QW i Q ,KW i K ,VW i V ) (5) In the above formula, W i Q ,W i K ,W i V ∈ is the parameter of linear projection, n is the number of attention heads, and after Add&Norm, it is LayerNorm(x+Linerlayer(x)), where Norm is the smooth input feature; Position encoding is used in self-attention; then the calculated x is input into the MLP layer and the position convolutional encoding layer, and its formula is shown as follows; PE(Q,K,V)=Norm(x+MLP(x)) (6) Among them, MLP is a fully connected feedforward network that transforms the original features into arbitrary dimensions. Formula (7) is a position convolutional coding layer that encodes the image to retain the detailed information of the local features of the target.
Citation Information
Patent Citations
Behavior recognition method and system based on space attention and grouping convolution
CN114783053A
Multi-target tracking method and system based on random channel adaptive attention mechanism
CN114842388A
Video action recognition method based on multi-dimensional feature excitation network
CN115862137A
Captive panda action recognition method based on space-time channel attention mechanism
CN116844228A
Behavior recognition method, electronic device and computer-readable storage medium
WO2024120125A1