A Visual Object Tracking Method Based on Multi-Level Aggregation and Attention Siamese Network
Through multi-level aggregation and attention to twin network methods, integrating different levels of features, suppressing noise and enhancing semantic representation, the goal matching problem of twin network trackers in complex scenarios is solved, and more efficient visual target tracking is achieved.
Patent Information
- Application Number
- CN202010653926.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-07-08
AI Technical Summary
The existing visual target tracking method based on twin networks has the problem of degradation of target matching similarity when dealing with complex scenarios such as occlusion, deformation and background clutter changes, resulting in a reduced tracking performance.
The multi-level aggregation and attention twin network method is adopted to extract multi-layer features through the twin backbone network, combine high-level semantic features and low-level detail features, add self-refinement modules to suppress noise, and add header attention modules to enhance semantic representation at the top-level convolutional features to build a multi-level aggregation and attention twin network tracker.
It improves the accuracy and robustness of visual target tracking, enhances the ability to identify targets, and significantly improves tracking performance.
Smart Images

Figure CN111860249B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual object tracking, and in particular to a visual object tracking method based on multi-level aggregation and attention Siamese network. Background Art
[0002] Visual object tracking refers to automatically locating a specified target in a continuously changing video sequence. It is one of the most fundamental research issues in the field of computer vision and has extensive applications in visual surveillance, human-computer interaction, video editing, etc. The core problem of object tracking is how to accurately and effectively detect and locate the target in challenging scenarios with occlusion, out-of-view, deformation, and background clutter changes.
[0003] In recent years, trackers based on Siamese networks have shown great potential in visual tracking in terms of speed and robustness by converting the tracking problem into a similarity learning problem. During the offline training phase of the network, they use convolutional neural networks as the backbone network to learn features for classification or regression on the external large-scale video dataset ILSVRC2015. Different from handcrafted features, these backbone networks can not only generate well-organized feature representations but also have the generalization ability across datasets. Therefore, the tracker only needs to be trained offline and does not require any online fine-tuning of the network during the tracking process to ensure robust tracking, which is very gratifying. However, although the design of Siamese network trackers is convincing, they still inevitably have some limitations. Most tracking methods only use deep features, and usually, this feature representation has a low resolution, which may lead to the loss of some target-specific details and local structure information. Therefore, these trackers are often not sensitive enough to details and it is difficult to distinguish two targets with the same attributes or semantics. Summary of the Invention
[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the specification of this application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by the present invention is: to propose a visual object tracking method based on multi-level aggregation and attention Siamese network to solve the problem that the introduction of position deviation in the Siamese tracking framework reduces the matching similarity between the target and the search sample, thereby leading to a decrease in tracking performance.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: A visual object tracking method based on multi-level aggregation and attention Siamese network, comprising the following steps: using a Siamese backbone network to be responsible for extracting multi-layer feature representations of example samples and search samples; defining a multi-layer aggregation module, selectively integrating high-level semantic features and low-level detail features to learn complementary information between multi-layer features, so as to assist the shallow-layer features in tracking the target; adding a self-refinement module after the multi-layer aggregation module to suppress the noise generated by multi-layer aggregation; adding a head attention module at the top convolutional features of the Siamese backbone network to enhance the semantic representation of the top-level features and improve the recognition ability of the target; constructing a multi-level aggregation and attention Siamese network tracker for visual object tracking.
[0008] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network of the present invention, wherein: the Siamese backbone network comprises the following construction steps: adopting an improved ResNet22; dividing the Siamese backbone network into 3 stages, which includes 22 convolutional layers with a stride of 8; when the convolutional layer uses padding, using a cropping operation to eliminate the feature calculation affected by zero-padding, and keeping the internal block structure unchanged; following the original ResNet to perform feature downsampling in the first 2 stages of the network; in the 3rd stage, using max pooling with a stride of 2 to replace the convolutional layer to perform downsampling, and this layer is located in the first block of this stage, that is, layer2-1.
[0009] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network of the present invention, wherein: the Siamese backbone network includes two identical branches, namely an example branch and a search branch; wherein the example branch receives the input of the example sample; the search branch receives the input of the search sample; the two branches share parameters in the convolutional neural network to ensure the same transformation is used for these two samples; using the output features of the last 3 blocks of the 3rd stage of the ResNet22 network, that is, layer2-2, layer2-3 and layer2-4.
[0010] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network of the present invention, wherein: the multi-level aggregation module comprises the following steps: extracting the representations of three layers of features of the example sample generated on the Siamese backbone network as F z1 , F z2 and F z3 ; using the method of transposed convolution to sample the last 2 layers of features to the same resolution as F′ z2 and F′ z3 ; concatenating the three layers of features together, and performing a convolution operation on the concatenated features to generate the aggregated multi-level aggregation feature F M = conv(concat(Fz1 , F' z2 , F' z3 )), the F M fully encodes low-level detail information from the shallower layers and high-level semantic information from the deeper layers.
[0011] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network according to the present invention, wherein: after the multi-level aggregation module, adding a self-refinement module includes combining the representation of the multi-level aggregation features with the shallow feature F z1 and inputting them into the self-refinement module to generate the following refined features: where SrM(·) represents the self-refinement module; calculating the matching similarity by using the refined features and the shallow feature F corresponding to the search sample x1 ; the similarity calculation can be expressed as: where Corr(·) represents the cross-correlation operation.
[0012] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network according to the present invention, wherein: the self-refinement module includes using the aggregated features F z1 and F M as the input, and dividing the self-refinement module into two parts; in the first part, global average pooling is used in the channel direction of the input features to compress the feature spatial dependence, and then a 1×1 convolution conv z1 and the Sigmoid function σ are used to generate a channel mask u ∈ R 1×1 , and finally multiplying it with the input features, and the specific process is described as: C×1×1 where GAP is global average pooling,
[0013]
[0014]
[0015] represents element-wise multiplication, and F' represents the output feature of the first part.
[0016] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network according to the present invention, wherein: the self-refinement module includes, in the second part, using the output of the first part as the input; using a 3×3 convolution conv 3×3 to compress the input features, and then using the Sigmoid function σ for normalization operation to generate a spatial mask m ∈ R W×H×1 , and finally multiplying it with the input features, and the calculation process is expressed as:
[0017] m = σ(conv3×3 (F′)),
[0018]
[0019] where F″ is the final refined feature.
[0020] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network according to the present invention, wherein: the head attention module includes a spatial attention mechanism and a channel attention mechanism, and the spatial attention mechanism includes representing the input feature of the attention module as F ∈ R C×W×H , where C, W, and H represent the channel, width, and height dimensions respectively; inputting this input feature into 3 convolutional layers with the same structure to obtain 3 new features, namely F q , F k and F v , all belonging to R C×W×H ; reconstructing F q and F k into R C×N , where N = H × W; then performing matrix multiplication between F q and the transpose of F k , and applying the Softmax operation to generate a spatial attention map:
[0021]
[0022] Define F sji to represent the influence of the feature at position i on the feature at position j, and the closer the connection between the two, the larger the value of F sji ; reconstruct F v into R C×N , and perform matrix multiplication with F s to obtain the result as F r , multiply the F r by a parameter λ s , and perform an element-wise summation operation with F to obtain the final output as:
[0023]
[0024] where the value of λ s is initialized to 0 and gradually learns to assign more weights to the spatial attention map.
[0025] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network according to the present invention, wherein: the channel attention mechanism includes transforming the input feature F ∈ R C×W×H into R C×N, where N = H × W; then perform matrix multiplication on the transpose of F and F, and apply the Softmax operation to obtain the channel attention map F c ∈R N×N :
[0026]
[0027] where F cji is used to measure the influence of the i-th channel on the j-th channel. Similar to the spatial attention mechanism, the larger the value of F cji , the greater the mutual connection between the two; then reconstruct it into R C×N The feature F and the channel attention map F c Perform matrix multiplication, and multiply with the parameter λ c , and finally perform an element-wise summation operation with F to obtain the final output:
[0028]
[0029] where λ c is similar to that in spatial attention, initialized to 0 and gradually learned to control the channel importance of the input feature F.
[0030] As a preferred solution of the visual object tracking method based on multi-level aggregation and attention Siamese network of the present invention, wherein: after the spatial attention mechanism and the channel attention mechanism, the following steps are included. The two newly generated attention features perform an element-wise operation to obtain the spatial-channel attention feature F sca ; In the tracking framework SiamMLAA of the proposed multi-level aggregation and attention Siamese network, F is F z3 , similar to the shallow-layer similarity calculation, the deep-layer feature similarity calculation can be expressed as:
[0031]
[0032] where is the spatial-channel attention feature obtained by inputting F z3 into the head attention module.
[0033] Beneficial effects of the present invention: The multi-layer aggregation module is used to fully integrate features at different levels, the aggregated features are used to assist the shallow-layer features, and the self-refinement module is introduced to suppress noise and refine the features for better matching similarity calculation; in addition, a channel and spatial attention mechanism is added to the top layer of the Siamese backbone network to enhance the semantic information of the deep-layer features, resulting in a more significant improvement in the visual object tracking results. Brief Description of the Drawings
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. Among them:
[0035] Figure 1 It is a schematic diagram of the overall framework of the multi-level aggregation and attention Siamese network described in the present invention;
[0036] Figure 2 It is a schematic diagram of the overall process of the visual object tracking method based on the multi-level aggregation and attention Siamese network described in the present invention;
[0037] Figure 3 It is a schematic diagram of the structure of the multi-layer aggregation module described in the present invention;
[0038] Figure 4 It is a schematic diagram of the structure of the self-refinement module described in the present invention;
[0039] Figure 5 It is a schematic diagram of the structure of the head attention module described in the present invention;
[0040] Figure 6 It is a schematic diagram of the success graph and accuracy graph on OTB2013 described in the present invention;
[0041] Figure 7 It is a schematic diagram of the success graph and accuracy graph on OTB2015 described in the present invention;
[0042] Figure 8 It is a schematic diagram of the success graph and accuracy graph of the ablation experiment on OTB2013 described in the present invention;
[0043] Figure 9 It is the success graph and accuracy graph of the ablation experiment on OTB2015 described in the present invention. Detailed implementation manners
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Persons skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0046] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other.
[0047] The present invention is described in detail in conjunction with schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views showing the device structure are enlarged locally not in accordance with the general scale, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.
[0048] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationships indicated by terms such as "upper, lower, inner, and outer" are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0049] Unless otherwise clearly defined and limited in the present invention, the terms "mounted, connected, and coupled" shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may also be a mechanical connection, an electrical connection, or a direct connection, or may be indirectly connected through an intermediate medium, or may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0050] Embodiment 1
[0051] Refer to Figure 1The schematic diagram shows the overall framework of the multi-level aggregation and attention Siamese network proposed in this embodiment. Most existing trackers rely on the output features of the last layer of the Siamese backbone network to track the target, often ignoring the characteristics of features at different levels. Therefore, this embodiment proposes a new network, called the Siamese Multi-Level Aggregation and Attention Network (SiamMLAA), which includes a Head Attention (HA) module, a Multi-Level Aggregation (MLA) module, and a Self-Refinement (SR) module. The simple process can be described as adding the head attention module to the top convolutional layer of the backbone network to improve the feature representation, and modeling the broader and richer context of the top-level features by using spatial and channel attention; in addition, the multi-level aggregation module can effectively integrate low-level spatial features and high-level semantic features to assist the shallow features in calculating the matching similarity, and the subsequent self-refinement module further refines and enhances the input features.
[0052] The Siamese network tracking model SiamMLAA based on multi-level aggregation and attention proposed in this embodiment includes a Siamese backbone network and three additional modules, namely a multi-level aggregation module, a self-refinement module, and a head attention module.
[0053] Since it is noted that the deep feature representation of the convolutional neural network (CNN) has a high semantic level and can effectively distinguish targets of different categories; while the shallow feature representation has a high resolution and can capture rich structural detail information. This is very useful for accurate positioning and can also handle different targets with the same semantics in the same category well. Therefore, this embodiment designs a multi-level aggregation module to selectively integrate high-level semantic features and low-level detail features to learn the complementary information between multi-level features, and then assist the shallow features in tracking the target. At the same time, to suppress the noise generated by multi-level fusion, a self-refinement module is introduced after the multi-level aggregation module. Finally, a head attention module is added to the top layer of the backbone network to model the broader and richer context of the deep features, enhance the feature representation of specific semantics, and have stronger robustness to changes in the target appearance.
[0054] Therefore, the simple process of this embodiment is as follows: First, a Siamese network tracking model with multi-level aggregation and attention is proposed to calculate the target similarity at multiple levels to achieve target tracking. This model includes a multi-level aggregation module, a self-refinement module, and a head attention module; second, the multi-level aggregation module selectively integrates low-level detail information and high-level semantic information to assist in calculating the shallow target similarity. In addition, a self-refinement module is introduced to suppress the noise generated by the fusion; an attention module is added at the top convolutional features to enhance the semantic representation of the top-level features to improve the recognition ability of the target.
[0055] Refer to Figure 2The schematic diagram shows the overall process of the visual object tracking method based on multi-level aggregation and attention Siamese network in this embodiment. The shallow features from the convolutional neural network have high resolution and can capture rich detailed information, while the deep features have low resolution and high semantic levels; the high-level semantic features can effectively identify different categories of targets and have strong robustness to the appearance changes of the targets, while the rich spatial details can accurately locate the targets and avoid confusion with similar objects; therefore, in order to make full use of the different characteristics of multi-level features and make the tracking more robust and accurate, a Siamese network tracking framework based on multi-level fusion and attention is proposed.
[0056] In this embodiment, the proposed tracking framework SiamMLAA will be introduced in more detail, where the Siamese backbone network is responsible for extracting multi-level feature representations of the example sample and the search sample; the multi-level aggregation module makes full use of the complementary information of the multi-level features generated by the backbone network to assist the shallow features in calculating the similarity; and the head attention module is used to enhance the semantic representation of the top-level features.
[0057] More specifically, a visual object tracking method based on multi-level aggregation and attention Siamese network includes the following steps:
[0058] S1: Use the Siamese backbone network to be responsible for extracting multi-level feature representations of the example sample and the search sample.
[0059] It should be noted in this step that powerful feature representations are crucial for accurate and robust visual object tracking, and in recent years, deep neural networks have been proven to be effective in Siamese network-based trackers, so they can be used in Siamese network-based trackers, such as VGGNet, ResNet, and MobileNet, etc. However, it is worth mentioning that Siamese network-based trackers are all based on the fully convolutional nature and are only applicable to the case where the backbone network used does not have padding operations. Although the original ResNet can learn very powerful feature representations, the network uses padding operations, which will introduce position biases in the Siamese tracking framework, resulting in a decrease in the matching similarity between the target and the search sample, and thus leading to a reduction in tracking performance.
[0060] Therefore, in order to solve the above problems, in the network of the tracker, an improved ResNet22 is adopted as the Siamese backbone network, and the specific improvements are as follows:
[0061] The Siamese backbone network is divided into 3 stages, which includes 22 convolutional layers with a stride of 8;
[0062] When padding is used in the convolutional layer, use the cropping operation to eliminate the feature calculation affected by zero-padding and keep the internal block structure unchanged;
[0063] Perform feature downsampling following the original ResNet in the first two stages of the network;
[0064] In the third stage, max pooling with a stride of 2 is used to replace the convolutional layer to perform downsampling. This layer is in the first block of this stage, namely layer2-1 (there are a total of 4 blocks in this stage, namely layer2-1, layer2-2, layer2-3, and layer2-4).
[0065] Optionally, in this step, the Siamese backbone network includes two identical branches, namely the exemplar branch and the search branch; the exemplar branch receives the input of the exemplar sample, and the search branch receives the input of the search sample; the two branches share parameters in the convolutional neural network to ensure the same transformation is used for these two samples; to calculate the matching similarity of multi-layer features, the output features of the last 3 blocks in the third stage of the ResNet22 network, namely layer2-2, layer2-3, and layer2-4, are used.
[0066] S2: Define a multi-layer aggregation module to selectively integrate high-level semantic features and low-level detail features to learn the complementary information between multi-layer features to assist the shallow features in tracking the target.
[0067] It should be noted that: noticing that the multi-layer similarity can improve the recognition ability of the Siamese network, so different from the existing Siamese network that calculates similarity based on the features of the last layer, calculate the target similarity from multiple levels to improve the robustness of the tracker. However, independently processing each layer of features, that is, directly using shallow and high-level features for target tracking, is often not so effective. Therefore, in this step, considering the internal connection between different levels of features, a multi-layer aggregation module is proposed to fuse multi-layer features together to assist the shallow features in learning more discriminative target features to calculate similarity. Refer to Figure 3 as shown (explanation of the multi-layer aggregation module), this is very effective for accurate and robust visual target tracking. The specific multi-layer aggregation module includes the following steps.
[0068] Extract the representations of the exemplar sample generating three-layer features F z1 , F z2 and F z3 on the Siamese backbone network; since the features of the above three levels have different spatial sizes, the last two layers of features are sampled to the same resolution of F′ z2 and F′ z3 by means of deconvolution;
[0069] Concatenate the three-layer features together and perform a convolutional operation on the concatenated features to generate the aggregated multi-layer aggregation feature F M = conv(concat(Fz1 ,F′ z2 ,F′ z3 )), F M Fully encode low-level detail information from shallow layers and high-level semantic information from deep layers.
[0070] Furthermore, in this step, adding a self-refinement module after the multi-layer aggregation module includes:
[0071] Combine the representation of multi-layer aggregate features with the shallow features F z1 Combined and input into the self-refinement module, the following refinement features are generated: Where SrM(·) represents the self-refinement module;
[0072] When F M Multi-layer fusion features and shallow features F z1 When combined and fed into the self-refinement module, F M The shallow and high-level complementary information in F can be a good aid z1 Obtain a powerful feature representation by combining the refined features with the shallow features F corresponding to the search sample x1 To calculate the matching similarity, which is very helpful for the final tracking performance, the similarity calculation can be expressed as: where Corr(·) represents the cross-correlation operation.
[0073] S3: A self-refinement module is added after the multi-layer aggregation module to suppress the noise generated by multi-layer aggregation.
[0074] In the multi-layer aggregation module, the complementary information between features at different levels is combined to obtain a comprehensive feature expression. Although this method is simple and intuitive, it still has certain defects. For example, due to the contradictions between different levels, some noise will inevitably be introduced, affecting the final tracking effect. Therefore, a self-refinement module is developed to further refine and enhance the fused feature representation.
[0075] Specific reference Figure 4 , which shows the overall structure of the self-refining module in this embodiment.
[0076] The self-refinement module in this step includes:
[0077] With feature F z1 and F M The aggregate feature F z1 As input, the self-refinement module is divided into two parts;
[0078] In the first part, global average pooling is used according to the channel direction of the input feature to compress the feature space dependency, followed by a 1×1 convolution conv 1×1and the Sigmoid function σ to generate the channel mask u ∈ R C×1×1 , and finally multiply it with the input feature, and the specific process is described as:
[0079]
[0080]
[0081] where GAP is global average pooling, denotes element-wise multiplication, and F′ denotes the output feature of the first part.
[0082] Furthermore, the self-refinement module includes
[0083] In the second part, take the output of the first part as the input;
[0084] Adopt a 3×3 convolution conv 3×3 compress the input feature, and then use the Sigmoid function σ for normalization operation to generate the spatial mask m ∈ R W×H×1 , and finally multiply it with the input feature, and the calculation process is expressed as:
[0085] m = σ(conv 3×3 (F′)),
[0086]
[0087] where F″ is the final refined feature.
[0088] S4: Add a head attention module at the top convolutional feature of the siamese backbone network to enhance the semantic representation of the top feature and improve the recognition ability of the target.
[0089] Refer to Figure 5 for the schematic diagram, which shows the structural schematic diagram of the head attention module. As mentioned above, the shallow feature contains the spatial structure information of the target and can well locate the target. However, the discrimination ability of the network mainly comes from the semantic information of the deep feature. Therefore, it is particularly important to obtain strong semantic features. For this reason, an attention module is added to the last layer convolution of the siamese backbone network, and the regions more relevant to the semantic description of the target are emphasized together through the spatial and channel self-attention mechanisms, so as to obtain a stronger semantic feature representation.
[0090] In this step, the middle attention module includes a spatial attention mechanism and a channel attention mechanism, where the spatial attention mechanism includes
[0091] Express the input feature of the attention module as F ∈ R C×W×H , where C, W, and H represent the channel, width, and height dimensions respectively;
[0092] Input this input feature into three convolutional layers with the same structure to obtain three new features, namely F q , F k and F v , all belonging to R C×W×H ;
[0093] Reconstruct F q and F k into R C×N , where N = H × W;
[0094] After that, perform matrix multiplication between the transpose of F q and F k , and apply the Softmax operation to generate a spatial attention map:
[0095]
[0096] Define F sji to represent the influence of the feature at position i on the feature at position j. The closer the connection between the two, the larger the value of F sji ;
[0097] Reconstruct F v into R C×N , and perform matrix multiplication with F s to get the result as F r . Multiply F r by a parameter λ s , and perform an element-wise summation operation with F to obtain the final output as:
[0098]
[0099] where the value of λ s is initialized to 0 and gradually learns to assign more weights to the spatial attention map.
[0100] Furthermore, the channel attention mechanism includes,
[0101] Transform the input feature F ∈ R C×W×H into R C×N , where N = H × W;
[0102] After that, perform matrix multiplication between the transpose of F and F, and apply the Softmax operation to obtain the channel attention map F c ∈ R N×N :
[0103]
[0104] where F cji is used to measure the influence of the i-th channel on the j-th channel. Similar to the spatial attention mechanism, Fcji The larger the value, the greater the mutual connection between the two;
[0105] Then it is reconstructed into R C×N The feature F of and the channel attention map F c Perform matrix multiplication and multiply with the parameter λ c Multiply and finally perform an element-wise summation operation with F to obtain the final output:
[0106]
[0107] where λ c Similar to that in spatial attention, it is initialized to 0 and gradually learned to control the channel importance of the input feature F. From the final output of the above formula, it can be seen that the feature F at each position sa is the result of the weighted sum of all position features and the original feature. Therefore, it selectively aggregates the global context under the guidance of the spatial attention map, promotes each other for similar semantic features, and improves the intra-class compactness and semantic consistency.
[0108] It should be noted here that the channel attention mechanism: The spatial attention mechanism enhances the feature representation of specific semantics from a spatial perspective. In fact, different channel mappings of features can be regarded as responses to specific targets, and different semantic responses are interrelated. Therefore, by utilizing the correlation between channel mappings, the feature representation with specific semantics can also be enhanced. So it is very necessary for us to construct a channel attention module to establish the relationship between channels.
[0109] In this step, after the spatial attention mechanism and the channel attention mechanism, the following steps are included:
[0110] The two newly generated attention features perform an element-wise operation to obtain the spatial-channel attention feature F sca ;
[0111] In the tracking framework SiamMLAA of the proposed multi-level aggregation and attention Siamese network, F is F z3 , similar to the shallow similarity calculation, the deep feature similarity calculation can be expressed as:
[0112]
[0113] where is the spatial-channel attention feature obtained by inputting F z3 into the head attention module.
[0114] S5: Construct a multi-level aggregation and attention Siamese network tracker for visual object tracking.
[0115] Example 2
[0116] To verify the actual effect of the visual object tracking method based on multi-level aggregation and attention Siamese network proposed in the above embodiments, the experimental results of this embodiment on five common tracking benchmark datasets including OTB2013, OTB50, OTB2015, VOT2016, and VOT2017 show that this method is superior to the baseline tracker in various evaluation criteria and is also highly competitive among existing tracking methods. Therefore, the proposed network SiamMLAA has achieved very good performance in all aspects.
[0117] Specifically, the proposed network framework is implemented on PyTorch and trained using 4 GPUs on RTX2080Ti.
[0118] The training process is as follows: The backbone network and the rest are initialized using the ResNet22 model pre-trained on the ILSVRC classification dataset and random noise respectively, and offline training is performed on the object tracking dataset GOT10K. This dataset contains more than 10,000 video clips of moving objects in the real world, divided into more than 560 categories, and all the bounding boxes of the objects are manually marked, with a total of more than 1.5 million. The training set is preprocessed by dividing the images into example samples and search samples with sizes of 127×127 and 255×255 pixels respectively, and the entire framework is trained using the stochastic gradient descent (SGD) method with a momentum of 0.9 and a weight decay of 0.0005. The learning rate is set to 0.01 and exponentially decays to 0.00001 after 100 iteration cycles. The entire process is trained on four GPUs with a minimum batch size of 8.
[0119] The testing process is as follows: During the testing process, for each input image, if it is the initial frame, it is cropped and adjusted to a size of 127×127 using the given first frame label and used as an example sample to input into the network; if it is a subsequent frame, it is cropped and adjusted to a spatial size of 255×255 centered on the position tracked in the previous frame and used as a search sample to input into the network. After obtaining the final feature maps respectively, the similarity between the two is calculated through correlation operations to generate a 17×17 similarity map, and then the similarity map is upsampled using bicubic interpolation to obtain a more accurate localization.
[0120] Table 1: Performance comparison on 5 tracking benchmarks.
[0121]
[0122] The results of this method are evaluated on five public tracking benchmarks against some advanced methods, including OTB2013, OTB50, OTB2015, VOT2016, and VOT2017. For fair evaluation, the tracking results provided by the corresponding authors or the publicly available codes are used to retrain and adjust the training parameters to obtain their best tracking results for comparison with the method of the present invention. The specific results are shown in Table 1.
[0123] Evaluation on the OTB benchmark: Evaluation experiments are carried out on the public tracking benchmark datasets OTB2013, OTB50, and OTB2015, which contain 51, 50, and 100 fully annotated video sequences respectively. The success plot and precision plot of one-pass evaluation (OPE) are used to compare different trackers, such as Figure 6 and Figure 7 , from which the experimental results of the tracker SiamMLAA on OTB2013 and OTB2015 can be intuitively seen. Among them, in Figure 6 and Figure 7 , a is the schematic of the success plot and b is the schematic of the precision plot.
[0124] In addition, the SiamMLAA tracker is also compared with other trackers based on Siamese networks. The specific evaluation results are shown in Table 1. The experiments show that SiamMLAA has the best performance on the three OTB benchmark datasets, and the area under the curve (AUC) of its success plot reaches 0.705 / 0.648 / 0.674 respectively. Compared with the baseline tracker, improvements of 4.1% / - / 2.2% are obtained, which shows the superiority of this method. Finally, the SiamMLAA tracker is evaluated against some non-real-time advanced tracking methods, including CCOT and ECO, etc. All the results in Table 1 show that SiamMLAA of the present invention can achieve a good balance between running speed and tracking performance.
[0125] This embodiment is also evaluated on the VOT benchmark: The VOT challenge is the most important annual competition in the field of visual tracking. The two datasets VOT2016 and VOT2017 both consist of 60 video sequences and are designed to evaluate the short-term tracking performance of trackers.
[0126] They are used to test the proposed tracking method SiamMLAA. In the experiment, the expected average overlap (EAO) metric is used to evaluate the overall performance of the tracking algorithm, and both accuracy and robustness are considered. The tracker SiamMLAA is evaluated on the VOT2016 and VOT2017 benchmark datasets respectively and compared with some other trackers. In addition to the baseline SiamDW and advanced trackers in VOT2016 / 2017 (such as CCOT and ECO, etc.), it also includes some tracking methods based on Siamese networks (such as C-RPN and TADT, etc.). The specific experimental results of each metric for different trackers are also shown in Table 1. The EAO scores of SiamMLAA on VOT2016 / 2017 reach 0.387 / 0.298 respectively. Compared with the baseline SiamDW, the present invention obtains absolute gains of 5.2% and 3.2%, which shows that the tracking method of the present invention is still highly competitive compared with other trackers.
[0127] Embodiment 3
[0128] To verify the effectiveness of each key module designed in the proposed tracker, an ablation study is also conducted in this embodiment. The ablation experiment is carried out on the OTB benchmark, which includes three datasets: OTB2013, OTB50, and OTB2015. From Figure 8 and Figure 9 it can be intuitively found that among them, in Figure 8 and Figure 9 a is a schematic diagram of the success graph, and b is a schematic diagram of the precision graph. The tracker containing all modules (i.e., the multi-layer aggregation MLA module, the self-refinement SR module, and the head attention HA module) achieves almost the best tracking performance in terms of both precision and success rate, which proves that each module in the tracker proposed by the present invention is necessary and plays a very obvious improvement role in the final tracking performance.
[0129] Table 2: Ablation study of different component combinations on the OTB dataset.
[0130]
[0131] In the experiment, we used SiamDW as the baseline tracker and then added each module separately to illustrate the impact of each module on the tracking performance. As can be seen in detail from Table 2, after adding the multi-layer aggregation module, the AUC scores on the three datasets of OTB2013 / 50 / 2015 increased significantly from 0.663 / - / 0.652 of the baseline tracker to 0.682 / 0.623 / 0.667. When the head attention module was added only to the top-level features, the AUC scores also reached 0.677 / 0.620 / 0.656 compared with the baseline. The combination of these two modules had a more significant improvement on the tracking results. Finally, the self-refinement module was also added to the tracker and the best result required by the present invention was obtained.
[0132] In this embodiment, a Siamese tracking network with multi-layer fusion and attention (SiamMLAA) is proposed to implement the visual object tracking task. Considering the different characteristics of features at different levels, a simple and effective multi-layer aggregation module is designed to fully integrate features at different levels. Then we use the aggregated features to assist the shallow features and introduce a self-refinement module to suppress noise and refine the features for better matching similarity calculation. In addition, we also added a channel and spatial attention mechanism to the top layer of the Siamese backbone network to enhance the semantic information of the deep features. The experimental results on 5 public tracking benchmark datasets show that the performance of this network is very good.
[0133] It should be recognized that the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose the program is capable of running on a dedicated integrated circuit programmed for this purpose.
[0134] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.
[0135] Further, the method can be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when the storage medium or device is read by the computer, can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via wired or wireless networks. When such media include instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself. The computer program is capable of applying to input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0136] As used in this application, the terms "component", "module", "system", etc. are intended to refer to a computer-related entity, which can be hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to: a process running on a processor, a processor, an object, an executable file, a thread in execution, a program, and / or a computer. As an example, an application running on a computing device and the computing device can both be components. One or more components can exist in a process and / or thread in execution, and the components can be located in one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media having various data structures thereon. These components can communicate in a local and / or remote procedure manner through signals such as according to one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, and / or communicates with other systems via a network such as the Internet in a signal manner).
[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A visual object tracking method based on multi-level aggregation and attention Siamese network, characterized in that: including the following steps, using a Siamese backbone network to be responsible for extracting multi-layer feature representations of example samples and search samples; defining a multi-layer aggregation module to selectively integrate high-level semantic features and low-level detail features to learn complementary information between multi-layer features, so as to assist the shallow feature in tracking the target; adding a self-refinement module after the multi-layer aggregation module to suppress the noise generated by multi-layer aggregation; adding a head attention module at the top convolutional features of the Siamese backbone network to enhance the semantic representation of the top features and improve the recognition ability of the target; constructing a multi-level aggregation and attention Siamese network tracker for visual object tracking; the multi-layer aggregation module includes the following steps, Extract example samples to generate representations of three layers of features, namely F z1 , F z2 and F z3 on the twin backbone network; Use the deconvolution method to sample the features of the last two layers to the same resolution of F′ z2 and F′ z3 ; Cascade three layers of features together, and perform a convolution operation on the cascaded features to generate an aggregated multi-layer aggregated feature F M = conv(concat(F z1 , F′ z2 , F′ z3 ))), where the F M fully encodes low-level detailed information from the shallow layer and high-level semantic information from the deep layer; the head attention module includes a spatial attention mechanism and a channel attention mechanism, wherein the spatial attention mechanism includes, Represent the input feature of the head attention module as F ∈ R C×W×H , where C, W, and H represent the channel, width, and height dimensions respectively; Input this input feature into three convolutional layers with the same structure to obtain three new features, namely F q , F k and F v , all belonging to R C×W×H ; Refactor F q and F k into R C×N , where N = H × W; After that, at F q and the transpose of F k perform matrix multiplication and apply the Softmax operation to generate a spatial attention map: Define F sji Indicates the influence of the feature at position i on the feature at position j for measurement. The closer the connection between the two, the greater the value of F sji will be; Reconstruct F v into R C×N and perform matrix multiplication with F s to obtain the result F r Multiply the said F r by a parameter λ s and perform element-wise summation with F to obtain the final output as: where λ s is initialized to 0 and gradually learns to assign more weight to the spatial attention map; the channel attention mechanism includes, Transform the input feature F ∈ R C×W×H into R C×N , where N = H × W; Subsequently, perform matrix multiplication on the transpose of F and F, and apply the Softmax operation to obtain the channel attention map F c ∈R N×N : Among them, F cji is used to measure the influence of the i-th channel on the j-th channel. Similar to the spatial attention mechanism, F cji The larger the value, the greater the mutual connection between the two; Then reconstruct it into R C×N The feature F and the channel attention map F c Perform matrix multiplication and multiply with the parameter λ c Multiply, and finally perform an element-wise summation operation with F to obtain the final output: where λ c is similar to that in spatial attention, initialized to 0 and gradually learned to control the channel importance of the input feature F; after the spatial attention mechanism and the channel attention mechanism, it includes the following steps, The two newly generated attention features perform element-wise operations to obtain the spatial channel attention feature F sca ; In the proposed tracking framework SiamMLAA of the multi-level aggregation and attention Siamese network, F is F z3 , similar to the shallow similarity calculation, the deep feature similarity calculation can be expressed as: Among them is F z3 The spatial channel attention feature obtained by inputting into the head attention module.
2. The visual object tracking method based on multi-level aggregation and attention Siamese network according to claim 1, characterized in that: the Siamese backbone network includes the following construction steps, adopting an improved ResNet22; dividing the Siamese backbone network into 3 stages, which includes 22 convolutional layers with a stride of 8; when padding is used in the convolutional layer, use a cropping operation to eliminate the feature calculation affected by zero-padding and keep the internal block structure unchanged; perform feature downsampling following the original ResNet in the first 2 stages of the network; in the third stage, replace the convolutional layer with max pooling with a stride of 2 to perform downsampling, and this layer is located in the first block of this stage, that is, layer2-1.
3. The visual object tracking method based on multi-level aggregation and attention siamese network according to claim 1 or 2, characterized in that: the Siamese backbone network includes two identical branches, namely the example branch and the search branch; among them, the example branch receives the input of the example sample; the search branch receives the input of the search sample; the two branches share parameters in the convolutional neural network to ensure the same transformation is used for these two samples; use the output features of the last 3 blocks in the third stage of the ResNet22 network, that is, layer2-2, layer2-3, and layer2-4.
4. The visual object tracking method based on multi-level aggregation and attention Siamese network according to claim 3, characterized in that: adding a self-refinement module after the multi-layer aggregation module includes, Combine the representation of the multi-layer aggregation features with the shallow features F z1 and input them into the self-refinement module to generate the following refined features: where SrM(·) represents the self-refinement module; Combine the refined feature with the shallow feature F corresponding to the search sample x1 to calculate the matching similarity; The similarity calculation can be expressed as: where Corr(·) represents the cross-correlation operation.
5. The visual object tracking method based on multi-level aggregation and attention Siamese network according to any one of claims 4, characterized in that: the self-refinement module includes, With feature F z1 and F M as the aggregated feature as the input, the self-refining module is divided into two parts; In the first part, global average pooling is adopted in the channel direction of the input features to compress the feature space dependence, and then a 1×1 convolution conv 1×1 and the Sigmoid function σ are used to generate a channel mask u ∈ R C×1×1 , and finally it is multiplied by the input features. The specific process is described as follows: where GAP is global average pooling, represents element-wise multiplication, and F′ represents the output feature of the first part.
6. The visual object tracking method based on multi-level aggregation and attention Siamese network according to claim 5, characterized in that: the self-refinement module includes, in the second part, use the output of the first part as the input; Adopt a 3×3 convolution conv 3×3 Compress the input features and then perform a normalization operation using the Sigmoid function σ to generate a spatial mask m∈R W×H×1 , and finally multiply it with the input features. The calculation process is expressed as: m = σ(conv 3×3 (F′)), where F″ is the final refined feature.
Citation Information
Patent Citations
Sequence recommendation method based on self-attention network with deeper feature level
CN110083770A
Semantic similarity matching method based on twin network and multi-head attention mechanism
CN110781680A