A Visual Object Tracking Method APR-Net Based on Attention Pyramid Residual Network
By introducing attention pyramid residual network into the visual target tracking method, APR-Net solves the problem of insufficient tracking performance in the prior art in complex scenarios, achieving a balance between high precision and real-time performance.
Patent Information
- Application Number
- CN202310192160.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-02
AI Technical Summary
Existing visual target tracking methods are not effective when dealing with complex video sequences and target changes, especially in scenarios such as scale changes, lighting changes, occlusions and complex backgrounds, making it difficult to achieve a balance of high-precision and real-time performance.
A visual target tracking method APR-Net based on attention pyramid residual network is proposed. By introducing attention mechanisms and pyramid residual networks, robust appearance features are captured and attention modules are integrated in the tracking framework to improve the accuracy and robustness of feature extraction.
APR-Net significantly improves tracking performance, surpasses most advanced tracking methods, and achieves near real-time operation speeds, enabling accurate identification and tracking of targets in complex scenarios.
Smart Images

Figure CN116168061B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to target tracking technology, and specifically belongs to a visual target tracking method APR-Net based on attention pyramid residual network. Background Art
[0002] Visual object tracking is an indispensable research topic in computer vision. Due to its wide application, it has received extensive attention in various fields such as autonomous driving, video surveillance, motion analysis, and medical diagnosis [1-3]. Although visual object tracking methods have made considerable progress in the past few years, there is still much room for improvement compared to other computer vision tasks [3-8]. This is because, in most tracking tasks, the target is only known in the first frame of the video sequence, and the target is unknown in subsequent video frames. That is, in visual object tracking tasks, except for the first frame, the position, state and other information of the target object are unknown in advance in each subsequent frame of the video image.
[0003] It is relatively simple to identify and detect a specified object in a still image. However, it is difficult to robustly locate and track a specified target in continuous video images, and existing tracking methods often do not perform well. In long and complex video sequences that are not constrained by the environment, the video scene is constantly changing. More importantly, the target object may also experience various uncertain changes, such as significant deformation, interference from similar objects around it, scale changes, long or short occlusions or being out of the camera's field of view, lighting changes, and motion blur [3, 6, 7]. Under these severe challenges, not to mention computer vision tracking systems, even the human visual system, it is a huge challenge to continuously track a specified target. Therefore, accurately identifying the changes of the target and marking the target's bounding box in each frame is one of the milestone goals in the field of visual target tracking [9, 10].
[0004] Previous studies have shown that the appearance representation of the target directly affects the tracking performance of the tracking method. Therefore, building a robust feature representation model can meet the tracking system's needs for target changes during the inference phase. Generally speaking, according to different appearance representation methods, most tracking methods can be divided into two categories: tracking methods based on manual features and tracking methods based on deep features.
[0005] Before 2013, tracking methods based on hand-crafted features were one of the most widely used methods in the field of visual target tracking. The earliest tracking method based on hand-crafted features can be traced back to 1955, when Wax
[11] proposed a tracking method that used correlation between points to track radar signals. In 1960, Kalman
[12] proposed a linear Kalman filter for visual target tracking. In 1987, Sethi and Jain
[13] proposed a point-based visual target tracking method. Since then, tracking methods based on underlying features such as kernel histograms
[14] , mean shift
[15] , SIFT
[16] , mixture of Gaussians (MOGs)
[17] , histograms
[18] , histograms of oriented gradients (HOGs)
[19] , Hough transform
[20] , and particle filters
[21] have emerged one after another. Hand-crafted features require parameters to be manually designed based on experience, require a small number of parameters, and have a fast execution speed. However, manual features can hardly meet the requirements of tracking accuracy and robustness.
[0006] Since deep features were first successfully introduced into visual object tracking tasks in 2013
[22] , tracking methods have developed rapidly due to their successful application in visual object tracking. Deep features have greatly improved tracking accuracy and robustness due to their rich feature representation capabilities. At present, tracking methods based on deep features dominate the tracking field. Among the various deep models used for visual object tracking, the convolutional neural network (CNN) family is one of the most influential and widely used deep models [4]. However, the expensive computational complexity of convolutional features still hinders its application in practical application scenarios with high requirements for real-time tracking performance. Therefore, in visual tracking tasks based on deep models, how to improve computational efficiency without losing accuracy or making a trade-off between tracking accuracy and real-time performance is one of the important issues that need to be studied.
[0007] The goal of the present invention is to capture robust appearance features by introducing a pyramid residual network (PyConvResNet)
[23] with an attention mechanism
[24] in the tracking framework, thereby achieving a competitive tracking method called APR-Net. Standard convolutional networks are based on kernels of a single spatial size and the same depth, so their multi-scale processing capabilities are limited. Unlike standard convolution, APR-Net contains kernels of different sizes and depths, so objects of different sizes can be automatically captured through different convolutional layers. Compared with other existing network architectures, APR-Net is more efficient in terms of the number of parameters and the amount of computation. Therefore, the feature extraction model APR-Net proposed in the present invention is more suitable for tracking tasks.
[0008] The tracking method proposed in this paper has achieved significant improvement in tracking performance, surpassing most advanced tracking methods and achieving near real-time running speed. At the same time, it is proved that the quality of feature representation depends largely on the full end-to-end training of the feature extraction model. After sufficient training, the feature extraction model based on the attention mechanism and PyConvResNet network architecture proposed in this paper can maximize the recognition of target objects from a bunch of similar objects in each frame.
[0009] [1] W. Luo, P. Sun, F. Zhong, W. Liu, T. Zhang, Y. Wang, End-to-end active object tracking and its real-world deployment via reinforcement learning, IEEE transactions on pattern analysis and machine intelligence 42(6) (2019) 1317–1332.
[0010] [2]J.Shen,Y.Liu,X.Dong,X.Lu,FSKhan,S.Hoi,Distilled siamese networks for visualtracking,IEEE Transactions on Pattern Analysis and MachineIntelligence 44(12)(2021)
[0011] 8896–8909.
[0012] [3]Y.Wu,J.Lim,M.-H.Yang,Object tracking benchmark,IEEE Transactionson Pattern Analysisand Machine Intelligence 37(9)(2015)1834–1848.
[0013] [4]H.Kiani Galoogahi,A.Fagg,C.Huang,D.Ramanan,S.Lucey,Need for speed:A benchmarkfor higher frame rate object tracking,in:Proceedings of the IEEEInternational Conference onComputer Vision,2017,pp.1125–1134.
[0014] [5]K.Cannons,A review of visual tracking,Dept.Comput.Sci.Eng.,YorkUniv.,Toronto,
[0015] Canada,Tech.Rep.CSE-2008-07 242(2008).
[0016] [6]L.Huang,X.Zhao,K.Huang,Got-10k:A large high-diversity benchmarkfor generic objecttracking in the wild,IEEE Transactions on Pattern Analysisand Machine Intelligence 43(5)
[0017] (2019)1562–1577.
[0018] [7]M.Mueller,N.Smith,B.Ghanem,A benchmark and simulator for uavtracking,in:Europeanconference on computer vision,Springer,2016,pp.445–461.
[0019] [8]H.Yang,L.Shao,F.Zheng,L.Wang,Z.Song,Recent advances and trends invisual tracking:
[0020] A review,Neurocomputing 74(18)(2011)3823–3831.
[0021] [9]H.Fan,L.Lin,F.Yang,P.Chu,G.Deng,S.Yu,H.Bai,Y.Xu,C.Liao,H.Ling,Lasot:Ahigh-quality benchmark for large-scale single object tracking,in:Proceedings of theIEEE / CVF conference on computer vision and patternrecognition,2019,pp.5374–5383.
[0022]
[10] M.Muller,A.Bibi,S.Giancola,S.Alsubaihi,B.Ghanem,Trackingnet:Alarge-scale datasetand benchmark for object tracking in the wild,in:Proceedings of the European conference oncomputer vision(ECCV),2018,pp.300–317.
[0023]
[11] N.Wax,Signal-to-noise improvement and the statistics of trackpopulations,Journal ofApplied physics 26(5)(1955)586–595.
[0024]
[12] R.E.Kalman,et al.,A new approach to linear filtering andprediction problems[j],Journal ofbasic Engineering 82(1)(1960)35–45.
[0025]
[13] I.K.Sethi,R.Jain,Finding trajectories of feature points in amonocular image sequence,IEEETransactions on Pattern Analysis and MachineIntelligence 1(PAMI-9)(1987)56–73.
[0026]
[14] D.Comaniciu,V.Ramesh,P.Meer,Real-time tracking of non-rigidobjects using mean shift,in:Proceedings IEEE Conference on Computer Visionand Pattern Recognition.CVPR 2000
[0027] (Cat.No.PR00662),Vol.2,IEEE,2000,pp.142–149.
[0028]
[15] D.Comaniciu,V.Ramesh,P.Meer,Kernel-based object tracking,IEEETransactions onpattern analysis and machine intelligence 25(5)(2003)564–577.
[0029]
[16] D.G.Lowe,Object recognition from local scale-invariant features,in:Proceedings of theseventh IEEE international conference on computervision,Vol.2,Ieee,1999,pp.1150–1157.
[17] C.Stauffer,W.E.L.Grimson,Learningpatterns of activity using real-time tracking,IEEETransactions on patternanalysis and machine intelligence 22(8)(2000)747–757.
[0030]
[18] F.Porikli,Integral histogram:A fast way to extract histograms incartesian spaces,in:2005
[0031] IEEE Computer Society Conference on Computer Vision and PatternRecognition(CVPR’05),
[0032] Vol.1,IEEE,2005,pp.829–836.
[0033]
[19] Z.Yin,F.Porikli,R.T.Collins,Likelihood map fusion for visualobject tracking,in:2008
[0034] IEEE Workshop on Applications of Computer Vision,IEEE,2008,pp.1–7.2
[0035]
[20] A.Goldenshluger,A.Zeevi,The hough transform estimator,The Annalsof Statistics 32(5)(2004)1908–1932.
[0036]
[21] M.Isard,A.Blake,Contour tracking by stochastic propagation ofconditional density,in:
[0037] European conference on computer vision,Springer,1996,pp.343–356.
[0038]
[22] N.Wang,D.-Y.Yeung,Learning a deep compact image representationfor visual tracking,
[0039] Advances in neural information processing systems 26(2013).
[0040]
[23] I.C.Duta,L.Liu,F.Zhu,L.Shao,Pyramidal convolution:rethinkingconvolutional neuralnetworks for visual recognition,arXiv preprint arXiv:2006.11538(2020).
[0041]
[24] Q.Wang,Z.Teng,J.Xing,J.Gao,W.Hu,S.Maybank,Learning attentions:residualattentional siamese network for high performance online visualtracking,in:Proceedings of theIEEE conference on computer vision and patternrecognition,2018,pp.4854–4863.
[0042]
[25] M.Kristan,A.Leonardis,J.Matas,M.Felsberg,R.Pflugfelder,L.ˇCehovinZajc,T.Vojir,G.
[0043] Bhat,A.Lukezic,A.Eldesokey,et al.,The sixth visual object trackingvot2018 challengeresults,in:Proceedings of the European Conference onComputer Vision(ECCV)Workshops,
[0044] 2018,pp.0–0.
[0045]
[26] J.Fan,W.Xu,Y.Wu,Y.Gong,Human tracking using convolutional neuralnetworks,IEEEtransactions on Neural Networks 21(10)(2010)1610–1623.
[0046]
[27] M.Danelljan,G.Bhat,F.Shahbaz Khan,M.Felsberg,Eco:Efficientconvolution operators fortracking,in:Proceedings of the IEEE conference oncomputer vision and pattern recognition,
[0047] 2017,pp.6638–6646.
[0048]
[28] G.Bhat,J.Johnander,M.Danelljan,F.S.Khan,M.Felsberg,Unveiling thepower of deeptracking,in:Proceedings of the European Conference on ComputerVision(ECCV),2018,pp.
[0049] 483–498.
[0050]
[29] F.Zhang,S.Chang,Hierarchical convolutional features fusion forvisual tracking,in:Journalof Physics:Conference Series,Vol.1651,IOPPublishing,2020,p.012134.
[0051]
[30] F.Li,C.Tian,W.Zuo,L.Zhang,M.-H.Yang,Learning spatial-temporalregularizedcorrelation filters for visual tracking,in:Proceedings of the IEEEconference on computervision and pattern recognition,2018,pp.4904–4913.
[0052]
[31] Z.Zhu,Q.Wang,B.Li,W.Wu,J.Yan,W.Hu,Distractor-aware siamesenetworks for visualobject tracking,arXiv e-prints(2018)arXiv–1808.
[0053]
[32] T.Zhang,C.Xu,M.-H.Yang,Robust structural sparse tracking,IEEEtransactions on patternanalysis and machine intelligence 41(2)(2018)473–486.
[0054]
[33] H.Fan,H.Ling,Sanet:Structure-aware network for visual tracking,in:Proceedings of theIEEE conference on computer vision and patternrecognition workshops,2017,pp.42–49.
[0055]
[34] H.Nam,B.Han,Learning multi-domain convolutional neural networksfor visual tracking,in:
[0056] Proceedings of the IEEE conference on computer vision and patternrecognition,2016,pp.
[0057] 4293–4302.
[0058]
[35] L.Zhang,A.Gonzalez-Garcia,J.v.d.Weijer,M.Danelljan,F.S.Khan,Learning the modelupdate for siamese trackers,in:Proceedings of the IEEE / CVFinternational conference oncomputer vision,2019,pp.4010–4019.
[0059]
[36] Y.Zhang,L.Wang,J.Qi,D.Wang,M.Feng,H.Lu,Structured siamese networkfor real-timevisual tracking,in:Proceedings of the European conference oncomputer vision(ECCV),2018,
[0060] pp.351–366.
[0061]
[37] D.Held,S.Thrun,S.Savarese,Learning to track at 100 fps with deepregression networks,in:
[0062] European conference on computer vision,Springer,2016,pp.749–765.
[0063]
[38] L.Bertinetto,J.Valmadre,J.F.Henriques,A.Vedaldi,P.H.Torr,Fully-convolutionalsiamese networks for object tracking,in:European conference oncomputer vision,Springer,
[0064] 2016,pp.850–865.
[0065]
[39] K.O’Shea,R.Nash,An introduction to convolutional neural networks,arXiv preprintarXiv:1511.08458(2015).
[0066]
[40] A.Krizhevsky,I.Sutskever,G.E.Hinton,Imagenet classification withdeep convolutionalneural networks,Communications of the ACM 60(6)(2017)84–90.
[0067]
[41] K.He,X.Zhang,S.Ren,J.Sun,Deep residual learning for imagerecognition,in:Proceedingsof the IEEE conference on computer vision andpattern recognition,2016,pp.770–778.
[0068]
[42] F.Scarselli,M.Gori,A.C.Tsoi,M.Hagenbuchner,G.Monfardini,The graphneural networkmodel,IEEE transactions on neural networks 20(1)(2008)61–80.
[0069]
[43] M.Danelljan,G.Bhat,F.S.Khan,M.Felsberg,Atom:Accurate tracking byoverlapmaximization,in:Proceedings of the IEEE / CVF Conference on ComputerVision and PatternRecognition,2019,pp.4660–4669.
[0070]
[44] M.-H.Guo,T.-X.Xu,J.-J.Liu,Z.-N.Liu,P.-T.Jiang,T.-J.Mu,S.-H.Zhang,R.R.Martin,M.-M.Cheng,S.-M.Hu,Attention mechanisms in computer vision:Asurvey,ComputationalVisual Media(2022)1–38.
[0071]
[45] B.A.Olshausen,C.H.Anderson,D.C.Van Essen,A neurobiological modelof visualattention and invariant pattern recognition based on dynamic routingof information,Journal ofNeuroscience 13(11)(1993)4700–4719.
[0072]
[46] L.Chen,H.Zhang,J.Xiao,L.Nie,J.Shao,W.Liu,T.-S.Chua,Sca-cnn:Spatial andchannel-wise attention in convolutional networks for imagecaptioning,in:Proceedings of theIEEE conference on computer vision andpattern recognition,2017,pp.5659–5667.
[0073]
[47] A.Sagar,Dmsanet:Dual multi scale attention network,in:International Conference on ImageAnalysis and Processing,Springer,2022,pp.633–645.
[0074]
[48] S.Woo,J.Park,J.-Y.Lee,I.S.Kweon,Cbam:Convolutional blockattention module,in:
[0075] Proceedings of the European conference on computer vision(ECCV),2018,pp.3–19.
[0076]
[49] M.Yang,J.Yuan,Y.Wu,Spatial selection for attentional visualtracking,in:2007 IEEEConference on Computer Vision and Pattern Recognition,IEEE,2007,pp.1–8.
[0077]
[50] J.Fan,Y.Wu,S.Dai,Discriminative spatial attention for robusttracking,in:EuropeanConference on Computer Vision,Springer,2010,pp.480–493.
[0078]
[51] Z.Zhu,W.Wu,W.Zou,J.Yan,End-to-end flow correlation tracking withspatial-temporalattention,in:Proceedings of the IEEE conference on computervision and pattern recognition,
[0079] 2018,pp.548–557.
[0080]
[52] B.Huang,J.Chen,T.Xu,Y.Wang,S.Jiang,Y.Wang,L.Wang,J.Li,Siamsta:
[0081] Spatio-temporal attention based siamese tracker for tracking uavs,in:Proceedings of theIEEE / CVF International Conference on Computer Vision,2021,pp.1204–1212.
[0082]
[53] K.Yang,Z.He,Z.Zhou,N.Fan,Siamatt:Siamese attention network forvisual tracking,
[0083] Knowledge-based systems 203(2020)106079.
[0084]
[54] Y.Yu,Y.Xiong,W.Huang,M.R.Scott,Deformable siamese attentionnetworks for visualobject tracking,in:Proceedings of the IEEE / CVF conferenceon computer vision and patternrecognition,2020,pp.6728–6737.
[0085]
[55] J.Choi,H.J.Chang,J.Jeong,Y.Demiris,J.Y.Choi,Visual trackingusingattention-modulated disintegration and integration,in:Proceedings of theIEEE conference oncomputer vision and pattern recognition,2016,pp.4321–4330.
[0086]
[56] X.Chen,B.Yan,J.Zhu,D.Wang,X.Yang,H.Lu,Transformer tracking,in:Proceedings ofthe IEEE / CVF Conference on Computer Vision and PatternRecognition,2021,pp.
[0087] 8126–8135.
[0088]
[57] B.Jiang,R.Luo,J.Mao,T.Xiao,Y.Jiang,Acquisition of localizationconfidence for accurateobject detection,in:Proceedings of the Europeanconference on computer vision(ECCV),
[0089] 2018,pp.784–799.
[0090]
[58] B.Li,J.Yan,W.Wu,Z.Zhu,X.Hu,High performance visual tracking withsiamese regionproposal network,in:Proceedings of the IEEE conference oncomputer vision and patternrecognition,2018,pp.8971–8980.
[0091]
[59] B.Li,W.Wu,Q.Wang,F.Zhang,J.Xing,J.Yan,Siamrpn++:Evolution ofsiamese visualtracking with very deep networks,in:Proceedings of the IEEE / CVFConference on ComputerVision and Pattern Recognition,2019,pp.4282–4291.
[0092]
[60] T.Yang,P.Xu,R.Hu,H.Chai,A.B.Chan,Roam:Recurrently optimizingtracking model,in:
[0093] Proceedings of the IEEE / CVF conference on computer vision and patternrecognition,2020,
[0094] pp.6718–6727.
[0095]
[61] B.Yan,H.Zhao,D.Wang,H.Lu,X.Yang,’skimming-perusal’tracking:Aframework forreal-time and robust
[0096] long-term tracking,in:Proceedings of the IEEE / CVF InternationalConference on ComputerVision,2019,pp.2385–2393.
[0097]
[62] Z.Zhang,H.Peng,Deeper and wider siamese networks for real-timevisual tracking,in:
[0098] Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition,2019,
[0099] pp.4591–4600.
[0100]
[63] Y.Song,C.Ma,X.Wu,L.Gong,L.Bao,W.Zuo,C.Shen,R.W.Lau,M.-H.Yang,Vital:
[0101] Visual tracking via adversarial learning,in:Proceedings of the IEEEconference on computervision and pattern recognition,2018,pp.8990–8999.
[0102]
[64] K.Dai,D.Wang,H.Lu,C.Sun,J.Li,Visual tracking via adaptivespatially-regularizedcorrelation filters,in:Proceedings of the IEEE / CVFConference on Computer Vision andPattern Recognition,2019,pp.4670–4679.
[0103]
[65] M.Danelljan,G.Hager,F.Shahbaz Khan,M.Felsberg,Learning spatiallyregularizedcorrelation filters for visual tracking,in:Proceedings of the IEEEinternational conference oncomputer vision,2015,pp.4310–4318.
[0104]
[66] J.Valmadre,L.Bertinetto,J.Henriques,A.Vedaldi,P.H.Torr,End-to-endrepresentationlearning for correlation filter based tracking,in:Proceedingsof the IEEE conference oncomputer vision and pattern recognition,2017,pp.2805–2813.
[0105]
[67] A.Lukezic,T.Vojir,L.ˇCehovin Zajc,J.Matas,M.Kristan,Discriminative correlation filterwith channel and spatial reliability,in:Proceedings of the IEEE conference on computer visionand pattern recognition,2017,pp.6309–6318.
[0106]
[68] J.Zhang,S.Ma,S.Sclaroff,Meem:robust tracking via multiple expertsusing entropyminimization,in:European conference on computer vision,Springer,2014,pp.188–203.
[0107]
[69] Y.Li,J.Zhu,A scale adaptive kernel correlation filter trackerwith feature integration,in:
[0108] European conference on computer vision,Springer,2014,pp.254–265.
[0109]
[70] Z.Hong,Z.Chen,C.Wang,X.Mei,D.Prokhorov,D.Tao,Multi-store tracker(muster):Acognitive psychology inspired approach to object tracking,in:Proceedings of the IEEEconference on computer vision and pattern recognition,2015,pp.749–758.
[0110]
[71] T.Zhang,C.Xu,M.-H.Yang,Learning multi-task correlation particlefilters for visualtracking,IEEE transactions on pattern analysis and machineintelligence 41(2)(2018)
[0111] 365–378.
[0112]
[72] C.Ma,J.-B.Huang,X.Yang,M.-H.Yang,Robust visual tracking viahierarchicalconvolutional features,IEEE transactions on pattern analysis andmachine intelligence 41(11)
[0113] (2018)2709–2723.
[0114]
[73] L.Zhang,J.Varadarajan,P.Nagaratnam Suganthan,N.Ahuja,P.Moulin,Robust visualtracking using oblique random forests,in:Proceedings of the IEEEconference on computervision and pattern recognition,2017,pp.5589–5598.
[0115]
[74] A.He,C.Luo,X.Tian,W.Zeng,A twofold siamese network for real-timeobject tracking,in:
[0116] Proceedings of the IEEE conference on computer vision and patternrecognition,2018,pp.
[0117] 4834–4843.
[0118]
[75] M.Che,R.Wang,Y.Lu,Y.Li,H.Zhi,C.Xiong,Channel pruning for visualtracking,in:
[0119] Proceedings of the European Conference on Computer Vision(ECCV)Workshops,2018,pp.
[0120] 0–0.
[0121]
[76] J.Choi,H.J.Chang,T.Fischer,S.Yun,K.Lee,J.Jeong,Y.Demiris,J.Y.Choi,Context-aware deep feature compression for high-speed visualtracking,in:Proceedings of theIEEE conference on computer vision and patternrecognition,2018,pp.479–488.
[0122]
[77] Q.Wang,L.Zhang,L.Bertinetto,W.Hu,P.H.Torr,Fast online objecttracking andsegmentation:A unifying approach,in:Proceedings of the IEEE / CVFconference onComputer Vision and Pattern Recognition,2019,pp.1328–1338.
[0123]
[78] M.Danelljan,G.Hager,F.Khan,M.Felsberg,Accurate scale estimationfor robust visualtracking,in:British Machine Vision Conference,Nottingham,September 1-5,2014,BmvaPress,2014,p.1.
[0124]
[79] T.-Y.Lin, M.Maire, S.Belongie, J.Hays, P.Perona, D.Ramanan, P.Dollar, CLZitnick, Microsoft coco: Common objects in context, in: European conference on computer vision, Springer, 2014, pp.740–755. Summary of the invention
[0125] The present invention aims to solve the above problems of the prior art. A visual target tracking method APR-Net based on attention pyramid residual network is proposed. The technical solution of the present invention is as follows:
[0126] A visual target tracking method based on attention pyramid residual network APR-Net, which includes the following steps:
[0127] Design an attention pyramid residual network feature extraction model APR-Net, wherein the attention pyramid residual network feature extraction model APR-Net is a pyramid residual network with an attention mechanism;
[0128] The ATOM tracking method is improved, and a tracking framework is designed, which includes four parts: a feature extractor based on the pyramid network PyConvResNet, an attention module, a classifier, and a predictor; among them, the feature extractor based on the pyramid network PyConvResNet is used to extract multi-scale features of each frame of video image; the attention module is used to enhance the visual expression ability of the features, which helps the feature extractor pay more attention to the areas where the target may appear. The attention module is integrated into the pyramid network model through bit-by-bit weighted operations, and finally the tracking problem is converted into a classification and prediction problem; the classifier is used to preliminarily locate the target and obtain the initial target border, and the predictor uses back propagation to optimize the target border, and the tracking result of each frame is corrected to obtain a fine target border.
[0129] Furthermore, the pyramid residual network PyConvResNet extracts multi-scale convolution features using filters of different sizes and depths in each convolution layer; PyConvResNet inputs an image of a fixed size and automatically learns the multi-scale features in each convolution layer through filters of different sizes and depths.
[0130] Furthermore, the feature extraction model APR-Net consists of two parts: a pyramid feature layer and a hybrid attention module; the pyramid network is composed of four pyramid residual blocks, and a hybrid attention module is respectively placed after the third and fourth pyramid residual blocks; the hybrid attention layer is processed as a bitwise weighted operation; the final feature map generated by the feature extraction model APR-Net contains various scale features and attention features of the target.
[0131] Furthermore, the present invention replaces ResNet-18 in the original ATOM with APR-Net to extract more robust target appearance features; then an attention module is placed between PyConvResNet and the pooling layer of the region of interest PrRoIPooling.
[0132] The advantages and beneficial effects of the present invention are as follows:
[0133] First of all, the present invention integrates the visual attention mechanism into the PyConvResNet network architecture and specially designs a more robust and more suitable model for the tracking task, which the present invention names the attention pyramid residual network feature extraction model APR-Net. This model helps the feature extractor pay more attention to the target and know the most likely position where the target appears. This is crucial for identifying small objects, size changes of objects, deformations, out-of-view and viewpoint changes. APR-Net is a more effective alternative to the original deep residual network architecture.
[0134] Secondly, it is found through experiments that sufficient end-to-end training of the feature extraction model APR-Net on multiple datasets is very helpful for improving the tracking performance. Thus, it can be seen that the training strategy is one of the key factors for improving the tracking performance.
[0135] Thirdly, the present invention successfully applies the proposed APR-Net to a unified tracking framework. The tracking method of the present invention is a simple and effective tracking framework, which makes a good balance among tracking accuracy, robustness and tracking speed.
[0136] Finally, detailed experiments and analyses are carried out on the tracking method proposed by the present invention on multiple benchmark datasets. The experimental results show that by introducing the feature extraction model APR-Net and fully training on multiple datasets including LaSOT [1], TrackingNet
[29] and got10k
[30] , the tracking method designed by the present invention is comparable to or superior to the existing tracking methods in terms of tracking accuracy, robustness and real-time ability. Brief Description of the Drawings
[0137] Figure 1 It is the overall structure of the tracking framework proposed by the present invention.
[0138] Figure 2 The original and ranked graphs are the accuracy / stability (A / R) of the tracking methods on the VOT-2018 challenge benchmark dataset.
[0139] Figure 3 Original and ranked graphs of the accuracy / robustness (A / R) of six different visual attributes (a) camera motion, (b) illumination change, (c) motion change, (d) occlusion, (e) scale change, and (f) null attributes on the VOT-2018 dataset. DETAILED DESCRIPTION
[0140] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.
[0141] The technical solution of the present invention to solve the above technical problems is:
[0142] This paper will focus on the progress of two types of research work closely related to the present invention and some representative tracking methods, namely, reviewing the progress of existing tracking methods based on deep learning (Section 2.1) and attention mechanism (Section 2.2).
[0143] 2.1 Tracking method based on deep features
[0144] Tracking methods based on deep features can be traced back to 2010, when Fan et al.
[26] proposed a tracking method specifically for people by training convolutional neural networks (CNNs) to learn spatial and temporal features in images of two consecutive video frames. Reference
[26] achieved good tracking performance using convolutional neural networks. However, the tracking method proposed in reference
[26] is limited to tracking people and cannot effectively track other general targets. 2013 was a significant year for tracking methods based on convolutional networks. Wang Qiang et al.
[22] first proposed a general tracking method based on a deep learning architecture. First, they used stacked denoising autoencoders to learn the general appearance representation of the target object offline. Then, the offline learned features were transferred to the online tracking process, automatically obtaining valuable features from the previous video images. The average tracking speed of this tracking method on a GPU is 15 frames per second. On this basis, a series of deep feature-based tracking methods have been proposed for general visual target tracking tasks, such as ECO
[27] , UPDT
[28] , HCF
[29] , DeepSTRCF
[30] , DaSiamPRN
[31] , RSSTDeep
[32] , SANet
[33] , MDNet
[34] , UpdateNet
[35] and StructSiam
[36] .
[0145] So far, although existing tracking methods can achieve robust tracking performance on benchmark datasets, their tracking performance still has a lot of room for improvement in challenging scenarios such as scale changes, illumination changes, fast motion, very small targets, complex backgrounds, out-of-plane rotations, in-plane rotations, occlusions, etc. In addition, compared with the successful application of deep features in image classification and detection tasks, the application of deep features in visual object tracking tasks faces greater challenges due to the higher requirements for real-time performance. Due to the limitation of computational cost, the computational efficiency of tracking methods based on deep features is quite low. This makes it difficult for these traditional deep convolutional feature tracking methods to meet practical applications with high real-time requirements. Most of the existing tracking methods based on deep convolutional networks require the help of GPUs to achieve real-time tracking performance.
[0146] Recently, a series of deep learning-based tracking methods have been proposed to solve the problem of low computational efficiency [27, 37, 38]. Although these tracking methods have achieved success in real-time performance, it is usually difficult to achieve high accuracy and robustness. For most deep tracking methods, the most important task is how to balance tracking accuracy, robustness, and real-time performance.
[0147] Inspired by the successful application of pyramid residual network in image classification and video action classification / recognition
[23] , this paper proposes a simple and efficient tracking method (named APR-Net tracking method) based on the combination of multi-scale features and attention mechanism. This paper uses pyramid residual network (PyConvResNet) as the backbone network and uses attention mechanism to enhance the expression ability of visual features.
[0148] Standard convolutional networks, such as CNNs
[39] , AlexNet
[40] , ResNet
[41] , and GNN
[42] , have a single type of filter kernels with the same spatial resolution and depth in different convolutional layers. Therefore, in order to handle the size variation of objects in video images, these deep tracking methods solve the multi-scale problem through a brute force search method. Unlike popular standard convolutional networks that apply filters of the same size and depth, PyConvResNet uses filters of different sizes and depths in the convolutional layers to extract multi-scale convolutional features. PyConvResNet takes a fixed-size image as input and can automatically learn multi-scale features in each convolutional layer through filters of different sizes, so it can easily capture multi-scale features in each video image. As introduced in the literature
[23] , due to the flexibility of the filters, PyConvResNet can gradually reduce the size of the feature map, so that object features of various scales can be captured in a series of image patches.
[0149] Compared with the benchmark tracking method ATOM
[43] , the tracking method proposed in the present invention can achieve more competitive tracking performance without increasing the computation and storage capacity. The APR-Net of the present invention effectively improves the real-time performance without affecting the tracking accuracy and robustness. The present invention fully considers the trade-off between tracking performance and tracking speed when designing the entire visual target tracking framework. The experimental results in Section 6 show that by introducing variable-sized filters in the convolutional layer, the tracking accuracy and speed can be effectively improved. This is because the variable-sized filters increase scale invariance and reduce overfitting and computational complexity.
[0150] 2.2 Tracking method based on attention mechanism
[0151] Attention is a unique mechanism of the biological visual nervous system. When people look at an image, they focus on the target area rather than the entire image. Therefore, it is also hoped that the feature extraction network can focus on the area where the target may appear, just like the human visual system. That is, by introducing the attention mechanism, the feature network can focus on the key parts of the image or video. The attention mechanism adaptively adjusts the weights of the features according to the importance of the features in the input image, thereby automatically finding the salient areas in various scenarios
[44] . In 1993, Olshausen et al.
[45] first applied the attention mechanism to the field of neuroscience and achieved good performance. This was a pioneering work in introducing the attention mechanism into pattern recognition. Since then, the attention mechanism has quickly expanded to other research areas of computer vision, such as image annotation
[46] , image classification
[47] , object detection
[48] , image saliency
[44] , etc.
[0152] With the development and maturity of the attention mechanism, it has also been successfully transplanted to the field of visual target tracking. Researchers' studies have shown that the attention mechanism can effectively improve tracking performance. Representative tracking methods based on the attention mechanism include references
[49] ,
[50] ,
[51] and references [52, 53, 24, 54]. These tracking methods are all based on a basic discovery that by introducing the attention mechanism, more attention can be paid to the area where the target may appear, thereby improving the ability to recognize the target in the field of visual tracking.
[0153] In 2016, Choi et al. proposed a tracking method called AtCF in the literature
[55] , which is a pioneering work that introduced the attention mechanism into the field of visual target tracking. In the literature
[49] , a new tracking method AVT based on the attention mechanism was proposed. This tracking method improves the tracking performance by introducing an attention module to enhance the discrimination of the tracking model. In the literature
[50] , a selective attention tracking paradigm was introduced in the visual appearance model to improve the tracking performance. In addition, a series of tracking methods have attempted to improve the tracking performance by integrating the attention mechanism into the twin network tracking framework, such as the tracking methods SiamSTA
[52] and SiamAttn
[53] . However, these tracking methods that integrate attention into the twin network framework have difficulty in achieving good performance in complex scenes such as occlusion, complex background, and out of field of view, and due to the high computational complexity, it is difficult to meet the real-time requirements.
[0154] Afterwards, many attention-based tracking methods have emerged, such as the tracking method based on cross attention and self-attention proposed in
[56] , and the tracking method based on spatial and temporal attention proposed in
[50] and
[51] . Compared with the baseline tracking method, the above tracking methods with added attention mechanism achieve better tracking performance on the benchmark tracking dataset. This is because the attention mechanism effectively enhances the expressiveness of features.
[0155] The main factors affecting target tracking performance are: the target's feature model, the quality of the classifier and regressor, and the prediction performance, which reflect the accuracy of target recognition and detection and the robustness of target positioning. Therefore, in order to better simulate the human visual system, the present invention carefully designs a unified tracking framework. The addition of an attention module in the feature map is to make the feature extraction network pay more attention to the area where the target may appear in the feature map. This helps to enhance the appearance representation of the target and obtain robust and accurate tracking performance.
[0156] Integrating the attention mechanism with the pyramid residual network hardly increases the complexity of the network model and realizes the fusion of multi-scale information in the feature map with the attention mechanism. Most importantly, the attention mechanism helps to capture global dependencies from the feature map and effectively construct long-term dependencies
[56] , which is very suitable for visual object tracking tasks. Compared with ATOM, by introducing APR-Net, a better balance can be achieved between tracking accuracy, robustness, and real-time performance.
[0157] 3. Tracking method proposed by the present invention
[0158] In this section, we first describe the proposed overall tracking framework in Section 3.1. The attention pyramid residual network architecture will be introduced in Section 3.2. Sections 3.3 and 3.4 will introduce the attention pyramid network structure and its tracking process. Finally, Section 3.5 will focus on the similarities and differences between the tracking method of the present invention and ATOM.
[0159] 3.1 Overview of the tracking method proposed in this invention
[0160] In the present invention, in order to overcome the limitation that ATOM cannot automatically adapt to multi-scale and complex scenes, inspired by the core ideas of PyConvResNet
[23] and hybrid attention mechanism
[24] , a better tracking method is implemented by designing a pyramid residual network that integrates a hybrid attention mechanism. On this basis, a simple and effective tracking framework is designed, which includes four parts: a feature extractor based on the pyramid residual network PyConvResNet, an attention module, a classifier and a predictor. In short, the present invention follows the idea of the ATOM tracking framework, and the proposed tracking method converts the tracking problem into a classification and prediction problem by integrating the attention module into the pyramid residual network model.
[0161] A classifier is used to distinguish the target from the background and obtain the initial target border; a predictor is used to correct the tracking result of each frame; and the prediction result is refined through a modulation module. However, unlike the previous ATOM, in this patent, the ResNet-18 in ATOM is first replaced by APR-Net to extract the target appearance features. Then a hybrid attention module is placed between the pyramid residual network PyConvResNet and the pooling layer of the region of interest (PrRoIPooling)
[57] . Through the above improvements, the tracking method proposed in the present invention can better capture global information through a larger receptive field and multi-scale information without increasing the computational cost. Figure 1 The overall tracking framework of the present invention is described.
[0162] Figure 1The overall structure of the tracking framework proposed by the present invention. The tracking framework of the present invention consists of four components: a feature extractor based on a pyramid residual network, an attention module, a classifier, and a predictor. When the image of the template frame and the image of the current frame flow into the feature extractor respectively, a feature representation is generated by the feature extractor. The feature map enhanced by the attention module is aggregated to a fixed size using a PrRoIPool layer. The features output by the image of the current frame, the predicted target bounding box obtained in the image of the current frame, the features output by the template image and its predicted target bounding box are simultaneously input into the IoU predictor. The pooled features with the attention mechanism are modulated by performing channel-by-channel multiplication operations with the coefficient vector returned by the template branch. The classifier generates a confidence value based on the features output by the image of the current frame to obtain an initial target bounding box. Finally, a refined predicted target bounding box is obtained by maximizing the IoU overlap rate.
[0163] 3.2 Attention Pyramid Network Structure
[0164] This paper designs a deep feature extraction model that is more suitable for visual target tracking tasks, which is named APR-Net. The overall network architecture of APR-Net and its attention mechanism will be introduced in Sections 3.2.1 and 3.2.2 respectively.
[0165] 3.2.1 Overall network architecture
[0166] In previous visual target tracking tasks, attention and pyramid networks were explored respectively. This paper successfully applies the attention mechanism to the pyramid residual network for the first time to extract the appearance features of the target object in the tracking task. This paper draws inspiration from the attention mechanism
[24] and the pyramid residual network (PyConvResNet)
[23] . The backbone architecture of the feature extraction network model of the present invention is the deep residual network with jump connections (ResNet-18) proposed by He Kaiming et al.
[41] . Different from the original ResNet-18, APR-Net integrates the attention module into PyConvResNet to extract features from each frame. The feature extraction model APR-Net has filters of different sizes and depths in each layer, which can automatically extract multi-scale feature representations from each feature map.
[0167] The original pyramid residual network PyConvResNet is designed for target detection, image classification and image recognition. In order to enable it to be applied to tracking tasks, in the present invention, the original PyConvResNet architecture is modified and an attention mechanism is added, so that the feature extraction model APR-Net can make full use of the complementary advantages of the pyramid architecture and the attention mechanism. With slight modifications, the pyramid residual network architecture is more suitable for visual target tracking tasks. Therefore, in the tracking process, the present invention uses multi-scale deep features with an attention mechanism for general feature description of the target.
[0168] As the name implies, the model APR-Net for extracting deep features consists of two parts: a pyramid feature layer and an attention module. The pyramid feature network consists of four pyramid residual blocks, and the third and fourth pyramid residual blocks are followed by a hybrid attention module. The present invention processes the hybrid attention layer as a bit-by-bit weighted operation. The final feature map generated by the feature extraction model APR-Net contains various scale features and attention features of the target, which will emphasize the areas where the target may appear and suppress its background areas. In this way, the sign extraction model APR-Net can focus on the important areas where the target is most likely to appear in each frame of the image, thereby better locating the specified target.
[0169] The APR-Net architecture enables the feature extraction process to capture various types of complementary information from local information to global information. These improvements make the feature extraction model APR-Net more robust to various complex scenes in each sequence, especially in complex backgrounds, scale changes, deformations, motion blur, fast motion, etc., the feature extraction model APR-Net can better identify targets.
[0170] In addition, the integration of attention modules with multi-scale feature maps and the sufficient training of network models on multiple datasets are the key to improving tracking performance. This is mainly because the APR-Net network architecture can not only encode objects of various scales from feature maps, but also emphasize areas where objects may appear, which makes the learned features tightly coupled with the tracking process.
[0171] 3.2.2 Attention Mechanism
[0172] The present invention applies a hybrid attention module to learn the interdependencies between different features. The main motivation for designing the attention module is to emphasize the area where the target is likely to appear and ignore its background area. In the present invention, the attention module is integrated into the classifier branch and the predictor branch respectively. The attention module in the classifier guides the feature extraction network to enhance the area where the target is most likely to appear in each frame, thereby enhancing the robustness to target deformation. The intuition behind the introduction of the attention module is that not all features provide the same contribution to the classifier and the predictor. The attention module in the predictor can fine-tune the result of the predicted border to obtain a more accurate target border.
[0173] As shown in the literature
[24] , the present invention introduces the attention factor λ into the feature x and learns the deep appearance features through the attention mechanism. The formula is:
[0174]
[0175] in, represents a bit-by-bit weighted operation. In order to reduce the computational complexity, inspired by the literature
[24] , the attention factor λ is further decomposed into dual attention ε and channel attention κ. Dual attention ε is used to enhance the focus on the target, and channel attention κ is used to define the feature channel. Therefore, the attention is decomposed into:
[0176]
[0177] Among them, ε can be further decomposed into superposition attention, which consists of general attention (similar to Gaussian distribution) and residual attention (can be regarded as a discriminator) composition:
[0178]
[0179] The superposition attention module can help the feature model capture global saliency information in various video scenes, enhance the feature model's ability to identify targets, and reduce computational complexity. The hybrid attention module described in the literature
[24] is used as the template feature vector, which combines three types of attention: spatial attention, superposition attention, and channel attention.
[0180] The most important thing about introducing mixed attention is that after specifying the target in the first frame, the target candidate area that needs to be paid close attention to can be found in the image of each subsequent continuous video frame of the video sequence. The channel attention κ can further enhance the adaptability of the feature model to target changes. More details about mixed attention and its solution process are consistent with the reference
[24] . However, unlike the reference
[24] that introduces the attention mechanism on the twin network, the present invention integrates the attention mechanism into the pyramid residual network PyConvResNet, from local information to global information, capturing complementary information with multi-scale features.
[0181] 3.3 Feature Learning and Feature Extraction
[0182] The present invention adds hybrid attention at the end of the third and fourth pyramid residual blocks. The attention module is used to weight the feature map. More specifically, the template feature x from the first frame 0 and the test feature x of the current frame t t are transferred to their respective APR-Net branches to obtain their feature maps with attention mechanisms at position i. and and The higher the values at positions i, the more likely they are the target positions. and Lower values at position i are more likely to correspond to the background region of the object.
[0183] Feature map generated by template image and its predicted target bounding box Image features of the current frame and its predicted target bounding box At the same time, it serves as the input of the IoU predictor branch network. The classification branch network calculates the target confidence score according to the feature map to obtain the initial target border. The tracking method of the present invention can obtain complementary information from different convolutional layers through different types of filter kernels, thereby providing better discrimination ability for the learning of deep features.
[0184] 3.4 Tracking process
[0185] The essential problem of the tracking method proposed in this invention is how to learn classifiers and predictors. Therefore, in this subsection, the following two aspects will be described in detail: target classification and target prediction. During the tracking process, the tracking method performs target classification and target prediction in an alternating cycle. The goal of this invention is to learn a feature extraction network that maps the current frame from an image to a feature map through a series of pyramid convolution operations. After hybrid attention enhancement, the multi-scale feature map with attention is fed into the classifier and predictor.
[0186] 3.4.1 Target Prediction
[0187] The prediction branch is used to combine the modulation vector generated previously, infer the IoU overlap rate of the target border, and optimize the target border by back propagation to maximize its IoU and obtain a fine predicted border. The input template image size of each frame is 288x288, and the size of all output feature maps of each layer is adjusted to 38x38. The template branch is used to generate the first frame image feature x 0 Model the appearance of the target. The test branch is used to extract the appearance features x of the target in the current frame t t , and calculate the confidence and IoU overlap values.
[0188] The output of the template branch is two 1χ1χD z The modulation vector The test branch obtains a size of 1x1xD through the PrRoI pooling operation z The eigenvector of Finally, the output of the test branch is the IoU overlap rate, and its formula is:
[0189]
[0190] Among them, x 0 and are the template features and the labeled bounding boxes in each template image. The modulation vector composed of positive coefficients and According to the template characteristics and the initial bounding box Pre-calculated. Obtain target-specific appearance information through the modulation module and refine the predicted IoU overlap rate. Specified target prediction bounding box yes The parameterized representation of (c m ,c n ) is the position of the center of a bounding box. The subscripts m and n are the coordinates of the image, and w and h represent the width and height of the bounding box, respectively. The backbone feature map x is passed through the attention module and predicted bounding box The generated size is An expression of Where K is the spatial output size. It is an IoU estimation module composed of three fully connected layers. According to the given annotated bounding box, the prediction error of equation (4) can be minimized. The predicted target state can be obtained by maximizing equation (4).
[0191] 3.4.2 Target Classification
[0192] The goal of the classification branch is to preliminarily locate the target in a continuous video sequence. Therefore, in order to capture the target in real time in the image of each frame of the video, the classifier is learned online. The target classification component in the tracking framework of the present invention is defined as:
[0193]
[0194] in, is the backbone feature map of the image, ω={ω 1 ,ω 2} are network parameters, and is the activation function, and * represents standard multi-channel convolution.
[0195] The objective function of the classification error, inspired by the discriminative correlation filter (DCF), can be defined as:
[0196]
[0197] Used to label each training sample feature map x t,j , is a Gaussian function centered at the target location. Weight α j For control The impact of is the weight ω k Regularization.
[0198] 3.4.3 Target Tracking
[0199] The present invention uses the ATOM tracking method as the benchmark tracking framework
[43] and defines the tracking framework as a classifier and a predictor. The present invention uses a classifier to distinguish the target from its background, obtains the initial target bounding box, and combines a predictor to fine-tune the tracking result. Finally, the classifier and the predictor are jointly trained to optimize the tracking result and obtain the final tracking result in the current frame t.
[0200] In the target localization stage, a rectangular box is used to give the tracking target in the first frame. For subsequent video frames, the tracking method determines the position of the target border in the image of the current frame based on the state information of the target in the previous frame. During the tracking process, all output feature maps of each layer of the head are used to predict and classify the tracking target of each video frame. First, the classification module is used to predict and classify the tracking target based on the position P of the previous frame t-1. t-1 and size S t-1 Calculate the confidence map, with the position P with the highest confidence score t As the position of the target in the current frame t. The position P at the current frame t tThe size S calculated in the previous frame t-1 t-1 The initial target frame Next, 10 candidate regions are generated and their IoU overlap ratios are calculated using the target prediction module. The bounding boxes corresponding to the first three maximum values are taken and their average value is used as the final tracking bounding box.
[0201] 3.5 Discussion
[0202] This section will focus on analyzing and discussing the similarities and differences between the tracking method proposed in this invention and ATOM in terms of multi-scale search (Section 3.5.1) and attention mechanism (Section 3.5.2).
[0203] 3.5.1 Benefiting from Multi-Scale Search
[0204] The size of the target changes in almost any video sequence, so the tracking performance depends largely on the ability of the tracking method to identify the change in target size. Therefore, one of the long-term goals of visual object tracking is to be able to process inputs at multiple scales to accurately capture moving targets in video sequences. ATOM attempts to automatically identify scale changes from targets by maximizing the IoU overlap rate of target estimation components. The potential problem of ATOM is the degradation of tracking performance due to the inconsistency between the predicted bounding box and the true bounding box. This essentially explains why the original ATOM cannot further improve tracking accuracy and robustness. Therefore, the feature extraction network needs a more effective strategy to search for multi-scale information of the target.
[0205] Compared with ATOM, the tracking method proposed in the present invention can automatically capture the scale changes of the target through APR-Net. APR-Net is not limited to approximating a set of fixed-size filter kernels, but can be tuned to many filter kernels of different sizes and depths in each convolutional layer. The multi-scale features of each pyramid layer are combined with the original scale features of the current frame as the final features to accurately capture targets of different scales. Finally, the multi-scale features are used as feature maps of predictors and classifiers to capture targets of different sizes in the images of video frames. Therefore, when APR-Net is used as a feature extractor, it can provide significant advantages for tracking tasks.
[0206] 3.5.2 Benefiting from the Attention Mechanism
[0207] Unlike ATOM, which directly performs the PrRoIPool operation on the feature map, the present invention places a hybrid attention module between the PrRoIPool operation and the feature map generated by APR-Net. By adding an attention mechanism to the network, the attention map obtained by the feature extraction network model can pay more attention to the key areas.
[0208] Among them, the feature map with the attention module in the predictor focuses on the target area, and the feature map with the attention module in the classifier is fine-tuned by paying attention to the difference between the target and its background. Although the work of the present invention is an extension of the ATOM tracking method, it has achieved more competitive tracking performance than ATOM in various challenging scenarios.
[0209] 4 Implementation Details
[0210] The network model APR-Net and the proposed tracking method are implemented using Python on a computer with Intel i9 2.60GHz, 32.0GB RAM and NVIDIA RTX A2000GPU. The average speed of the proposed tracking method is close to 30 frames per second.
[0211] The settings in the training process of the present invention follow PyConvResNet
[23] and ATOM
[43] . To meet our GPU throughput requirements, all training images are trained with a minimum batch size of 64. The images or image sequences used for training are cropped to a fixed size of 288×288. Stochastic gradient descent (SGD) with a momentum of 0.9 is applied, and the weight decay is set to 0.005 during the end-to-end training phase. The network model of the present invention is trained for 50 epochs with a learning rate of 10 -5 .
[0212] In the present invention, APR-Net is used as the feature extractor of the tracking method. APR-Net is randomly initialized and first trained from scratch on the MS-COCO dataset in an offline manner. Then the predictor branch is fully trained offline on three dedicated tracking training datasets (LaSOT, TrackingNet and got10k). The predictor is enabled to learn feature representations with multi-scale and attention mechanisms from general objects. The classifier is trained online to capture dedicated tracking target features, which is conducive to improving the generalization and discrimination capabilities of the tracking method for specific targets, which is similar to ATOM.
[0213] 5 Experimental verification and result analysis
[0214] In this section, we hope to verify the effectiveness of the proposed tracking method by conducting a series of verification experiments on multiple representative benchmark tracking datasets, including LaSOT[9], UAV123[7], UAV20L[7], VOT-2018
[25] , and NFS[4]. At the same time, an ablation study is performed to observe the impact of each component of the proposed tracking method on the tracking performance. The experiments in this section show that the APR-Net model can achieve more robust tracking performance without sacrificing computational efficiency.
[0215] 5.1 Experiment 1: LaSOT
[0216] As is known to all, the LaSOT dataset [9] consists of two types of datasets: one is a dedicated training subset that provides visual and language annotations to train deep tracking methods (1120 videos), and the other is a test subset that is used to evaluate the performance of tracking methods (280 videos). In this section, the LaSOT test subset is used to verify the effectiveness of the tracking method proposed in this paper. According to the evaluation protocol II in the literature [9], two indicators are used to evaluate the tracking performance: namely, tracking accuracy and success rate.
[0217] 5.1.1 Overall performance of LaSOT
[0218] The overall tracking performance of the proposed tracking method achieved a tracking accuracy of 0.524 and a success rate of 0.530 on the LaSOT validation subset. The proposed tracking method achieved better performance than the baseline tracking framework ATOM. Compared with ATOM, the tracking accuracy was improved by 1.9% and the success rate was improved by 1.6%.
[0219] The excellent performance can be attributed to the fact that the ResNet-18 network used by the ATOM tracking method to extract features is replaced by the proposed feature extraction model APR-Net. The reason for the success of the proposed tracking method on various tracking sequences is the design of the attention pyramid residual network, which gives the proposed tracking method a global receptive field. Therefore, it is very obvious to what extent the attention and pyramid networks contribute to this excellent tracking performance.
[0220] 5.1.2 Performance based on various properties of LaSOT
[0221] As described in the literature [9], LaSOT is more challenging than other mainstream benchmark datasets under the attributes of target scale change, leaving the camera field of view, and occlusion. Even so, the tracking method of the present invention still achieves a tracking accuracy of 0.521 and a success rate of 0.529 in the scale change attribute, and a tracking accuracy of 0.439 and a success rate of 0.448 in the out-of-field attribute. In addition, the tracking method of the present invention achieves a tracking accuracy of 0.476 and a success rate of 0.446 in the full occlusion attribute.
[0222] Most importantly, in addition to achieving competitive tracking performance on typical most challenging visual attributes (such as scale change, full occlusion and out-of-view), the tracking method proposed in the present invention has better tracking stability in various other challenging attributes (including viewpoint change, rotation, partial occlusion, motion blur, low resolution, illumination change, fast motion, deformation, camera motion, complex background and aspect ratio change). This shows that the tracking method of the present invention can adapt to various complex scenes and achieve satisfactory tracking results.
[0223] 5.2 Experiment 2: UAV123
[0224] The unmanned aerial vehicles (UAVs) dataset [7] consists of 123 high-definition low-altitude aerial viewpoint videos taken by professional drones, so it is also called the UAV123 benchmark tracking dataset. The evaluation indicators in UAV123 are the same as LaSOT, namely accuracy and success rate. Compared with other target tracking datasets in field scenes, tracking a very small target, a fast-moving target, a target outside the field of view, and a target in a complex and changing scene in the UAV123 scene is more challenging. Although the videos on the UAV123 benchmark dataset are full of huge challenges, the tracking method of the present invention still has strong tracking performance. The tracking accuracy of the tracking method proposed in the present invention is 0.854 and the success rate is 0.645. Therefore, it can be inferred that the tracking method of the present invention is very suitable for tracking targets in aerial videos in complex and changing real scenes (such as: camera perspective, height, position changes, low resolution, significant scale / aspect ratio changes, fast motion, very small objects, similar objects, and long-term partial or complete occlusion). The tracking method proposed in the present invention is significantly better than the baseline tracking framework ATOM tracking method of the present invention (tracking accuracy is 0.828, success rate is 0.621). Compared with ATOM, the tracking accuracy and success rate of the present invention are increased by 2.6% and 2.4% respectively.
[0225] This is mainly because the tracking method of the present invention can adapt well to changes in the target scale by extracting multi-scale features through the pyramid network. In addition, the integration of the attention module into the pyramid network facilitates the capture of features from different receptive fields and combines related features at different scales, so that the tracking method of the present invention can better identify various changes. Most importantly, the introduction of the attention mechanism in the pyramid network can allocate greater attention to the area where the target is located. Therefore, the tracking method of the present invention can stably track very small objects.
[0226] 5.3 Experiment 3: UAV20L
[0227] UAV123 contains three subsets, among which UAV20L is a subset of UAV123, including 20 long low-altitude high-definition videos. Compared with other short-term videos in UAV123, UAV20L is more challenging because it is easier to lose the target in the long video. The tracking method of the present invention is better than the ATOM tracking method, improving the tracking accuracy from 0.678 to 0.767 and the success rate from 0.522 to 0.585 on UAV20L. However, compared with the tracking performance of UAV123, the tracking performance of the tracking method of the present invention on UAV20L is slightly reduced (the tracking accuracy is reduced by 8.7% and the success rate is reduced by 6%). The tracking method of the present invention improves the target recognition ability by introducing a hybrid attention module into the original pyramid network to improve feature representation and tracking performance. Most importantly, the study in the literature
[29] found that the shallow layer can capture details including texture, edge, contour and color because the shallow layer has a small receptive field. Since the deep layer has a larger receptive field, more semantic information can be obtained. Therefore, feeding the 3rd and 4th residual blocks into their respective convolutional layers helps to fuse shallow features with deep features. Finally, the features of the 3rd and 4th residual blocks are used for classification and prediction, making full use of shallow detail information and high-level semantic information. These may be the key factors for the tracking method of the present invention to achieve ideal tracking performance even in long videos like UAV20L.
[0228] 5.4 Experiment 4: VOT-2018
[0229] As we all know, Visual Object Tracking (VOT) is a challenging benchmark dataset. Since 2013, ICCV or ECCV has reported improvements on this benchmark dataset and the results of challenging tracking methods almost every year. However, unlike other benchmark datasets for evaluation, accuracy / stability graphs are used to measure tracking performance on the VOT dataset. In this section, the experimental verification of the present invention is based on the standard of VOT-2018 (see http: / / votchallenge.net / vot2018).
[0230] The tracking method of the present invention is compared with well-known tracking methods such as ATOM
[43] , SiamDW
[62] , ECO
[27] , SA_Siam_P
[74] , SiamRPN
[58] , CPT
[75] , UpdateNet
[35] , TRACA
[76] , UPDT
[28] and SiamMask
[77] on the VOT-2018 dataset. Figure 2 As shown, the tracking method proposed in the present invention is located in the upper right corner of the VOT-2018 benchmark test, and its tracking performance is comparable to that of existing advanced tracking methods.
[0231] Figure 2 Original and ranked plots of the accuracy / stability (A / R) of tracking methods on the VOT-2018 challenge benchmark dataset. The tracking methods used for evaluation are ranked based on their average performance across all sequences. The top right tracking method is the best performing according to the VOT-2018 expected average overlap value.
[0232] In addition to the overall performance, Table 1 further reports the expected average overlap (EAO), overlap rate (Overlap), failure rate (Failures) and speed (frames / second) of the tracking methods. The tracking method of the present invention outperforms recent tracking methods in terms of expected average overlap rate and overlap rate, and is second only to UPDT in terms of stability (tracking failure rate). It is worth noting that the tracking method of the present invention runs at 28.2 frames / second, which is faster than the benchmark tracking framework ATOM (26.6 frames / second) of the present invention. The experimental results show that although the feature extraction network of ATOM is implemented based on the ResNet-18 backbone network, the feature extraction network APR-Net of the present invention, even if it is implemented based on the PyConvResNet-18 backbone network of the hybrid attention module, has a tracking speed comparable to that of ATOM.
[0233] Table 1 Tracking performance of tracking methods on the VOT-2018 benchmark dataset in terms of expected average overlap (EAO), overlap rate (Overlap), failure rate (Failures) and speed (frames / second).
[0234]
[0235] also, Figure 3 The tracking performance of the tracking method based on the attributes is shown. The tracking method proposed in the present invention is more suitable for scenes such as camera motion, illumination change, motion change, occlusion, scale change, and empty attributes. The tracking method of the present invention can effectively improve the tracking accuracy and robustness of VOT-2018. This is because the tracking method proposed in the present invention takes advantage of the attention mechanism to better locate the target in the global scope of the video. Therefore, the tracking method proposed in the present invention can meet the tracking performance requirements in these challenging attributes.
[0236] Figure 3 Accuracy / Robustness (A / R) rankings for 6 different visual attributes (A) camera motion, (B) illumination change, (C) motion change, (D) occlusion, (E) scale change, and (F) empty attributes on the VOT-2018 dataset. Tracking methods are ranked based on their average performance across all sequences. The top right tracking method is the best performing based on the expected average overlap value of VOT-2018.
[0237] 5.5 Experiment 5: Ablation Study on the NFS Dataset
[0238] Unlike low frame rate (30fps) video tracking benchmark datasets, NFS (Need for Speed) [4] is the first high frame rate video dataset and target tracking benchmark dataset. As described in [Need for Speed], the benchmark dataset consists of 100 sequences (380 frames) captured from real scenes by high frame rate (240PFS) cameras. The tracking performance on this dataset is ranked based on accuracy and real-time performance. Three variants of the tracking method proposed in this patent are implemented and evaluated in Section 5.5.1. In addition, the tracking performance of the feature extraction model APR-Net trained on different training datasets is reported in Section 5.5.2. It is hoped that these ablation experiment results can verify why the proposed tracking method can very effectively improve the tracking performance.
[0239] 5.5.1 Three variants of the tracking method proposed in this invention
[0240] In order to show how different components of the tracking method or its variants affect the final tracking performance, a comprehensive ablation study was performed on the NFS benchmark dataset. By transforming the ResNet-18 architecture, the impact of the feature extraction network APR-Net on the tracking performance was verified, based on three different variants of the model of the present invention, corresponding to the PyConvNet+ResNet model with only the pyramid layer, the Attention+ResNet model with only the attention layer, and the PyConvNet+Attention+ResNet model with both the pyramid layer and the attention layer. On the NFS benchmark dataset, under the same parameter settings, the tracking performance of the three variants of PyConvNet+ResNet, Attention+ResNet, PyConvNet+Attention+ResNet and ATOM was verified. The ATOM tracking method is used as the baseline tracking method of the present invention, and no modifications are made to the author's code in the comparison.
[0241] First, the attention module is simply removed from the network architecture and replaced with the original pyramid residual network
[23] based on ResNet-18 (PyConvNet+ResNet) without any other changes. After this change, the network structure for feature extraction consists only of a series of pyramid residual blocks without attention layers. Without considering hybrid attention, based only on PyConvNet+ResNet, the achieved tracking accuracy is 0.719 and the success rate is 0.599.
[0242] In addition to the pyramid residual network model helping to improve tracking performance, ablation experiments also found that the attention module has a certain impact on tracking performance. Compared with ATOM, when only the attention module is added to the original ResNet-18 (without the pyramid residual block), the tracking performance is slightly better (tracking accuracy is 0.692 and success rate is 0.581). The reason for achieving this tracking performance is that the model's discriminative ability is closely related to the introduction of the attention module in the model. It can be seen that the attention mechanism helps to enhance the ability of deep models to identify specific targets. It shows that the attention layer can effectively improve the tracking performance, which shows that the attention layer is very necessary in visual target tracking.
[0243] When the PyConvNet+ResNet backbone network is replaced with the proposed attention pyramid architecture PyConvNet+attention+ResNet, the tracking accuracy is improved by 2.1% and the success rate is improved by 1.4%. The attention pyramid network PyConvNet+attention+ResNet is also superior to the tracking method with only the attention module (without pyramid convolution) attention+ResNet. Therefore, the excellent performance of the proposed tracking method is attributed to the effective combination of PyConvNet+ResNet and the attention mechanism. It can be found from the results of the ablation study described above that in terms of tracking performance, none of the variants can exceed the tracking performance of the present invention that integrates attention into the pyramid network. In general, the feature extraction model with both attention and pyramid layers is superior to the feature extraction model with only attention or only pyramid layers in tracking performance.
[0244] Ablation results verify the superior performance of the proposed tracking method. Integrating attention and pyramid convolutions with different filter kernel sizes and depths into each residual block can provide significant advantages. This provides a powerful feature model for the tracking task, which is one of the main drivers of the excellent tracking performance.
[0245] In addition, the three variants of the feature extraction model of the present invention can always provide better tracking performance than the ATOM tracking framework. Compared with ATOM, the tracking method using only the attention module (attention+ResNet), the tracking method using only the pyramid architecture (PyConvNet+ResNet), and the tracking success rate using both attention and pyramid architecture (PyConvNet+attention+ResNet) increased by 1%, 2.8%, and 4.2%, respectively. The tracking method based on the PyConvNet+Attention+ResNet model achieved the most significant performance improvement (nearly 4.2% improvement in success rate over ATOM), because the feature extraction model of the present invention can automatically identify changes in the size of the target. Therefore, introducing the attention mechanism in the deep feature model can effectively improve the tracking performance.
[0246] In practice, many tracking methods can only achieve excellent performance on low-speed video sequences. The main reason is that these tracking methods require a lot of memory resources and have high computational complexity, resulting in slow calculation speed for each frame. Therefore, these tracking methods cannot achieve good tracking performance in high-speed video sequences. The tracking method of the present invention adopts attention mechanism and pyramid design, so that it can accurately capture targets of various sizes and stably locate targets in any video sequence, and shows very strong performance on the NFS dataset with higher frame rate (tracking accuracy of 0.740 and success rate of 0.613). These ablation study results illustrate the effectiveness of the feature extraction model proposed in the present invention.
[0247] 5.5.2 Training on different datasets
[0248] In this section, an experiment is performed to verify the effect of training a deep feature model on tracking performance under different training datasets. At the same time, the results of the proposed tracking method based on the feature extraction model APR-Net after training on different datasets are reported. After training the deep feature extraction model of the present invention on different training datasets such as MS COCO
[28] , LaSOT[1], TrackingNet
[29] and got10k
[30] , the area under the curve score (AUC), overlap rate precision (OP 0.50 ,OP 0.75 ) and the regularized tracking accuracy (Norm.Prec.) are shown in Table 2.
[0249] Tracking tasks are different from target recognition and detection tasks. Pre-training the feature extraction model only on the MS COCO dataset will make the tracking performance very unsatisfactory. Therefore, in addition to training the deep feature extraction model of the present invention on the MS COCO dataset, it is also necessary to train the deep feature extraction model proposed by the present invention on a specific tracking video sequence. Without any changes in the tracking framework, the tracking method proposed by the present invention provides the best tracking performance after training on all four datasets (including MS COCO, LaSOT, TrackingNet and got10k). However, when the present invention trains the network for feature extraction on a single or two training datasets, suboptimal tracking performance is achieved.
[0250] From the experimental results in Table 2, it can be seen that under the same feature extraction network, the results obtained after the feature extraction model is fully trained on multiple training data sets have better tracking performance than the results without sufficient pre-training. It can be seen that under the same feature extraction model, the more sufficient the training data, the more stable the tracking performance. The tracking performance after training the feature extraction model on a variety of different data sets shows the necessity of sufficient end-to-end training and fine-tuning on multiple data sets. Therefore, it can be concluded that for tracking methods based on the same feature extraction model, sufficient training and fine-tuning using multiple larger data sets is the reason for achieving strong tracking performance. This means that by increasing the number of training samples, the robustness of tracking can be effectively improved, thereby making the tracking performance better. Most importantly, using multiple different data sets for effective end-to-end training allows the various components in the feature extraction model to work tightly coupled.
[0251] Table 2 reports the area under the curve (AUC), overlap rate precision (OP) and other performance of the tracking method based on the feature extraction model APR-Net after pre-training the deep feature extraction model of the present invention on different training datasets (including MS COCO, LaSOT, TrackingNet and got10k). 0.50 OP 0.75 ), precision (Precision) and regularized precision (Norm.Prec.) and other aspects of tracking performance
[0252]
[0253] 6 Conclusion
[0254] In the present invention, a real-time and effective tracking method is developed by combining the advantages of the pyramid residual network and the attention mechanism, which meets the requirements of tracking accuracy and robustness. Experiments on benchmark tracking datasets (including LaSOT, UAV123, UAV20L, VOT-2018 and NFS) show that the tracking method proposed in the present invention outperforms advanced tracking methods in terms of tracking performance. The present invention has been verified through multiple experiments that the network architecture used for feature extraction and its sufficient training will inevitably play a good role in the field of visual target tracking.
[0255] Although the tracking method proposed in the present invention has made great progress in tracking performance, there is still a lot of room for improvement in tracking accuracy and real-time tracking performance. So far, there is no completely satisfactory solution because the existing solutions either trade complexity for accuracy or trade accuracy for inefficient computation. In future tracking tasks, we can try to better understand the contribution of each component of the tracking framework. More specifically, we hope to continue to further improve the tracking performance in terms of real-time, accuracy, and success rate. By integrating lightweight networks into the tracking framework, a faster and more stable tracking method can be achieved. At the same time, we also hope to encourage more researchers to study which components in the tracking framework can more effectively improve tracking performance.
[0256] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions.
[0257] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0258] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A visual target tracking method based on attention pyramid residual network APR-Net, It is characterized in that The following steps are involved: Design an attention pyramid residual network APR-Net feature extraction model, wherein the feature extraction model APR-Net is a pyramid residual network with an attention mechanism; Taking the ATOM tracking method as the baseline tracking framework, the designed tracking method includes four parts: a feature extractor based on the pyramid residual network PyConvResNet, an attention module, a classifier and a predictor; among them, the feature extractor based on the pyramid residual network PyConvResNet is used to extract the features of each frame of video image; the attention module is used to enhance the visual expression ability of the features, which helps the feature extractor to pay more attention to the areas where the target may appear; the attention module is integrated into the pyramid residual network model through bit-by-bit weighted operations, and finally the tracking problem is converted into a classification and prediction problem; the classifier is used to preliminarily locate the target and obtain the initial target border, and the predictor uses back propagation to optimize the target border, and the tracking result of each frame is corrected to obtain a fine target border; The feature extraction model APR-Net consists of two parts: a pyramid feature layer and a hybrid attention module; the pyramid network consists of four pyramid residual blocks, and the third and fourth pyramid residual blocks are followed by a hybrid attention module respectively; the hybrid attention layer is processed as a bit-by-bit weighted operation; the final feature map generated by the model APR-Net contains various scale features and attention features of the target; Replace the ResNet-18 in the original ATOM with APR-Net to extract more robust target appearance features; then place an attention module between PyConvResNet and the pooling layer PrRoIPooling of the region of interest; The attention module specifically includes: Introduce the attention factor λ into the feature x and enhance the deep visual features through the attention mechanism; Attention module The formula is: Among them, ⊙ represents the bit-by-bit weighted operation on the feature x. In order to reduce the number of parameters, the attention factor λ is further decomposed into the dual attention ε to increase the attention to the target, and the channel attention κ defines the feature channel. Therefore, the attention model can be decomposed into: In order to effectively capture the common features of objects in images of different video frames and the differences between objects in images of different video frames, ε can be further decomposed into and residual attention Composition of superposition attention:
2. A visual target tracking method based on attention pyramid residual network APR-Net according to claim 1, It is characterized in that The pyramid residual network PyConvResNet extracts multi-scale convolution features using filters of different sizes and depths in each convolution layer; a fixed-size image is input to PyConvResNet, and multi-scale features in each convolution layer are automatically learned through filters of different sizes and depths.
3. A visual target tracking method based on attention pyramid residual network APR-Net according to claim 1, It is characterized in that The function of the predictor is to accurately predict the border of the target by maximizing the IoU overlap rate in each frame through back propagation; the input template image size of each frame is 288×288, and the size of all output feature maps of each layer is adjusted to 38×38; the template branch is used to predict the border of the target based on the first frame image x 0 Model the appearance of the target and obtain multi-scale features with attention The test branch is used to extract the image x in the current frame t t Target appearance features And calculate the confidence value and IoU overlap value; where the subscript i represents the i-th position in the feature map; The output of the template branch is two 1×1×D z The modulation vector The test branch obtains a 1×1×D z The eigenvector of Finally, the output of the test branch is the IoU overlap rate, and its formula is:
4. A visual target tracking method based on attention pyramid residual network APR-Net according to claim 3, It is characterized in that The classifier specifically includes: The goal of the classifier is to separate the target from its background in each frame of a continuous video sequence and obtain an initial target bounding box; therefore, in order to capture the target in real time in the image of each frame of the video, the classifier is online learning; the target classification component in the tracking framework is defined as: The objective function of the classification error can be defined as:
5. A visual target tracking method APR-Net based on an attention pyramid residual network according to claim 4, characterized in that The target tracking and positioning specifically includes: in the target positioning stage, a rectangular box is used to give the tracking target in the first frame; for subsequent video frames, the tracking method determines the border position of the target in the current frame based on the state information of the target in the previous frame; in the tracking process, all output feature maps of each layer of the head are used to predict and classify the tracking target of each video frame; first, the classification module is used to predict and classify the tracking target according to the position P of the previous frame t-1. t-1 and size S t-1 Calculate the confidence map, with the position P with the highest confidence score t As the position of the target in the current frame t; the position P at the current frame t t The size S calculated in the previous frame t-1 t-1 The initial target frame Next, 10 candidate regions are generated and their IoU overlap ratios are calculated using the target prediction module. The bounding boxes corresponding to the first three maximum values are taken and their average value is taken as the final tracking bounding box.