Object Tracking Method and System Integrating Significant Information and Multi-Granularity Context Features
Through a target tracking system that fuses significant information and multi-grained context features, the problem of twin neural networks losing targets under occlusion, deformation and rotation is solved, and a higher precision and stable target tracking is achieved.
Patent Information
- Application Number
- CN202111671961.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-31
AI Technical Summary
Twin neural network-based target tracking algorithms are prone to lose targets under occlusion, deformation and rotation.
A target tracking system that fuses significant information and multi-grained context features, including twin neural networks, multi-branch fusion modules, global context modules, attention map modules and deep correlation modules. Through these modules, the accuracy of template features and the connection between search features and template features are enhanced.
It effectively avoids the loss of targets under occlusion, deformation and rotation, and improves the stability and accuracy of target tracking.
Smart Images

Figure CN114332488B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically relates to an object tracking method and system that integrates salient information and multi-granularity context features. Background Art
[0002] Object tracking is one of the basic problems in the field of computer vision. Through the object tracking algorithm, the template features of the object in the initial video frame containing the object can be extracted, and the position of the object in the subsequent video frames can be continuously and stably tracked through the template features of the object.
[0003] The object tracking algorithm based on the Siamese neural network obtains high-precision network parameters and matching functions through offline training with a large amount of data sets, and has the advantages of high precision and strong effectiveness. However, the object tracking algorithm based on the Siamese neural network has problems such as insufficient extraction of template features and lack of connection between video frames during the process of tracking the object. When there are occlusions, deformations, and rotations during the process of tracking the object, the object may be lost. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an object tracking method and system that integrates salient information and multi-granularity context features, so as to avoid the situation of losing the object when there are occlusions, deformations, and rotations during the process of tracking the object.
[0005] The specific technical solutions are as follows:
[0006] In the first aspect of the implementation of the present invention, an object tracking system that integrates salient information and multi-granularity context features is first provided, including a Siamese neural network, a multi-branch fusion module, a global context module, an attention map module, a depth cross-correlation module, and an object position determination module, where:
[0007] The Siamese neural network is used to obtain a template image and a search image, extract multiple features of the template image as template branch features, and extract multiple features of the search image as search branch features; the template image contains the appearance information of the object to be tracked; the search image is an image containing the object;
[0008] The multi-branch fusion module is used to obtain the template features of the template image according to the template branch features;
[0009] The global context module is used to obtain the search features of the search image according to the search branch features;
[0010] The attention map module is used to obtain the attention map of the search features and the attention map of the template features according to the search features and the template features;
[0011] The depth cross - correlation module is used to perform depth cross - correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map;
[0012] The target position determination module is used to perform classification and regression operations on the score map to determine the position of the target in the search image.
[0013] In the second aspect of the implementation of the present invention, a target tracking method that fuses significant information and multi - granularity context features is provided. The method is characterized in that the method is applied to a Siamese neural network, and the method includes:
[0014] Obtain a template image and a search image, extract multiple features of the template image as template branch features, and extract multiple features of the search image as search branch features; the template image contains the appearance information of the target to be tracked; the search image is an image containing the target;
[0015] Obtain the template feature of the template image according to the template branch features;
[0016] Obtain the search feature of the search image according to the search branch features;
[0017] According to the search feature and the template feature, obtain the attention map of the search feature and the attention map of the template feature;
[0018] Perform depth cross - correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map;
[0019] Perform classification and regression operations on the score map to determine the position of the target in the search image.
[0020] Optionally, the Siamese neural network includes Siamese sub - neural networks. The obtaining of the template image and the search image, and the extraction of multiple features of the template image as template branch features and the extraction of multiple features of the search image as search branch features include:
[0021] Obtain the template image and the search image through the Siamese sub - neural networks; the size of the search image is larger than the size of the template image;
[0022] Input the template image into the ResNet50 network of the Siamese neural network to obtain the vector convolution operation feature, two - dimensional matrix convolution operation feature, three - dimensional matrix convolution operation feature, four - dimensional matrix convolution operation feature, and five - dimensional matrix convolution operation feature of the template image, which are f t1 、f t2 、f t3 、f t4 、ft5 , as the template branch feature;
[0023] Input the search image into the twin neural network ResNet50 network, and the vector convolution operation feature, two-dimensional matrix convolution operation feature, three-dimensional matrix convolution operation feature, four-dimensional matrix convolution operation feature, and five-dimensional matrix convolution operation feature of the search image are respectively f s1 、f s2 、f s3 、f s4 、f s5 , as the search branch feature.
[0024] Optionally, the twin neural network includes a multi-branch fusion module. The process of obtaining the template feature of the template image based on the template branch feature includes:
[0025] Compress the channels of the f t3 、f t4 、f t5 features of the template branch feature to obtain the features f n3 、f n4 、f n5 ;
[0026] Pass the f t2 feature of the template branch feature through the multi-branch fusion module to obtain the feature f n2 ;
[0027] Add f n3 、f n4 、f n5 to f s2 respectively and perform a central cropping operation to obtain the template feature F t3 、F t4 、F t5 of the template image.
[0028] Optionally, the twin neural network includes a global context module. The process of obtaining the search feature of the search image based on the search branch feature includes:
[0029] Compress the channels of the f s3 、f s4 、f s5 features of the search branch feature to obtain the features f m3 、f m4 、f m5 ;
[0030] Pass the f m3 、f m4 、f m5 features through the global context module to obtain the search features F s3 、Fs4 , F s5 .
[0031] Optionally, the siamese neural network includes an attention map module. Obtaining the attention map of the search feature and the attention map of the template feature according to the search feature and the template feature includes:
[0032] Input the template feature F t3 , F t4 , F t5 and the search feature F s3 , F s4 , F s5 into the self-attention module and the cross-attention module of the attention map module respectively, to obtain the attention map of the template feature and the attention map of the search feature
[0033] Optionally, the siamese neural network includes a depth cross-correlation module. Performing depth cross-correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map includes:
[0034] Through the depth cross-correlation module, perform depth cross-correlation operations on the attention map of the template feature and the attention map of the search feature respectively to obtain score maps φ3, φ4, φ5.
[0035] Optionally, the siamese neural network includes a target position determination module. Performing classification and regression operations on the score map to determine the position of the target in the search image includes:
[0036] Input the score maps φ3, φ4, φ5 into the classification branch and the regression branch of the target position determination module respectively;
[0037] Through the classification branch, perform convolutions on the score maps φ3, φ4, φ5 with a convolution kernel size of 1×1 and a stride of 1 to obtain features with 2K channels Multiply with the preset learnable weights respectively to obtain classification features The classification features include the foreground and background features of the target in the search image;
[0038] Through the regression branch, perform convolutions on the score maps φ3, φ4, φ5 with a convolution kernel size of 1×1 and a stride of 1 to obtain features with 4k channels Multiply with the preset learnable weights respectively to obtain regression features The regression feature includes the feature of the target;
[0039] Based on the classification feature and the regression feature determine the position of the target in the search picture.
[0040] In another aspect of the embodiments of the present invention, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0041] The memory is used to store a computer program;
[0042] When the processor is used to execute the program stored on the memory, it implements the target tracking method for fusing significant information and multi-granularity context features described in any one of the above.
[0043] In another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the target tracking method for fusing significant information and multi-granularity context features described in any one of the above.
[0044] In another aspect of the embodiments of the present invention, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the target tracking method for fusing significant information and multi-granularity context features described in any one of the above.
[0045] An object tracking system that fuses significant information and multi-granularity context features provided by an embodiment of the present invention. The system includes a twin neural network, a multi-branch fusion module, a global context module, an attention map module, a deep cross-correlation module, and an object position determination module. By using this system, a template image and a search image can be obtained through the twin neural network, multiple features of the template image are extracted as template branch features, and multiple features of the search image are extracted as search branch features; through the multi-branch fusion module, a template feature of the template image is obtained according to the template branch features; through the global context module, a search feature of the search image is obtained according to the search branch features; through the attention map module, an attention map of the search feature and an attention map of the template feature are obtained according to the search feature and the template feature; through the deep cross-correlation module, the attention map of the template feature and the attention map of the search feature are subjected to deep cross-correlation to obtain a score map; through the object position determination module, classification and regression operations are performed on the score map to determine the position of the object in the search image. Through the multi-branch fusion module, the accuracy of template feature extraction is enhanced, and through the attention map module, the connection between the search feature and the template feature is enriched. It avoids the situation of losing the object that may occur when there are occlusions, deformations, and rotations during the process of tracking the object. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings:
[0047] Figure 1 It is a flowchart of an object tracking method that fuses significant information and multi-granularity context features provided by an embodiment of the present invention;
[0048] Figure 2 It is a schematic flowchart of an object tracking method that fuses significant information and multi-granularity context features provided by an embodiment of the present invention;
[0049] Figure 3 It is a block diagram of the multi-branch fusion module provided by an embodiment of the present invention;
[0050] Figure 4 It is a block diagram of the global context module provided by an embodiment of the present invention;
[0051] Figure 5 It is a block diagram of the attention map module provided by an embodiment of the present invention;
[0052] Figure 6 It is a test chart of the accuracy and success rate of the object tracking method provided by an embodiment of the present invention;
[0053] Figure 7EAO (Expected Average Overlap) value test chart of the target tracking method provided by an embodiment of the present invention;
[0054] Figure 8 EAO value test chart of the target tracking method provided by an embodiment of the present invention for various situations;
[0055] Figure 9 Representative visual result chart of the target tracking method provided by an embodiment of the present invention;
[0056] Figure 10 Schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] In the prior art, for the target tracking algorithm based on the Siamese neural network, there are problems such as insufficient extraction of template features and lack of connection between video frames during the target tracking process. When there are occlusions, deformations, rotations, etc. during the target tracking process, the target may be lost.
[0059] To solve the above problems, an embodiment of the present invention provides a target tracking system that fuses significant information and multi-granularity context features. The target tracking system provided by the embodiment of the present invention may include:
[0060] Siamese neural network, configured to obtain a template image and a search image, extract multiple features of the template image as template branch features, and extract multiple features of the search image as search branch features; the template image contains the appearance information of the target to be tracked; the search image is an image containing the target;
[0061] Multi-branch fusion module, configured to obtain the template feature of the template image according to the template branch features;
[0062] Global context module, configured to obtain the search feature of the search image according to the search branch features;
[0063] Attention map module, configured to obtain the attention map of the search feature and the attention map of the template feature according to the search feature and the template feature;
[0064] A depth cross - correlation module for performing depth cross - correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map;
[0065] A target position determination module for classifying and regressing the obtained score map to determine the position of the target in the search image.
[0066] Based on the target tracking system provided by the embodiments of the present invention, through the multi - branch fusion module, the accuracy of template feature extraction can be enhanced, and through the attention map module, the connection between the search feature and the template feature can be enriched. It avoids the situation of losing the target when there are occlusions, deformations, rotations, etc. during the process of tracking the target.
[0067] See Figure 1 , Figure 1 is a flowchart of a target tracking method that fuses significant information and multi - granularity context features provided by the embodiments of the present invention, and is applied to a Siamese neural network. The method may include the following steps:
[0068] S101, Obtain a template image and a search image, extract multiple features of the template image as template - branch features, and extract multiple features of the search image as search - branch features.
[0069] S102, Obtain the template feature of the template image according to the template - branch features.
[0070] S103, Obtain the search feature of the search image according to the search - branch features.
[0071] S104, According to the search feature and the template feature, obtain the attention map of the search feature and the attention map of the template feature.
[0072] S105, Perform depth cross - correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map.
[0073] S106, Perform classification and regression operations on the score map to determine the position of the target in the search image.
[0074] The template image contains the appearance information of the target to be tracked; the search image is an image containing the target.
[0075] Based on the target tracking method that fuses significant information and multi - granularity context features provided by the embodiments of the present invention, through the multi - branch fusion module, the accuracy of template feature extraction can be enhanced, and through the attention map module, the connection between the search feature and the template feature can be enriched. It avoids the situation of losing the target when there are occlusions, deformations, rotations, etc. during the process of tracking the target.
[0076] See Figure 2 , Figure 2Schematic flow chart of an object tracking method that integrates significant information and multi-granularity context features provided by an embodiment of the present invention.
[0077] In one implementation, the size of the input template image (Templateimage) can be 127×127×3. Then, the width and height of the input template image are 127×127 pixel points, and the number of channels is 3. The size of the input search image (Searchimage) can be 255×255×3. Then, the width and height of the input search image are 255×255 pixel points, and the number of channels is 3. The template image and the search image are respectively input into the template branch and the search branch, and the two branches are ResNet50 networks with shared parameters. Finally, the matrix convolution operation features of 5 dimensions of Resnet50 in the template branch and the search branch are respectively f t1 、f t2 、f t3 、f t4 、f t5 and f s1 、f s2 、f s3 、f s4 、f s5 . f t1 、f t2 、f t3 、f t4 、f t5 The sizes are 61×61×64, 31×31×256, 15×15×512, 15×15×1024, 15×15×2048 respectively. f s1 、f s2 、f s3 、f s4 、f s5 The sizes are 125×125×64, 61×61×256, 31×31×512, 31×31×1024, 31×31×2048 respectively.
[0078] In one embodiment, the Siamese neural network includes Siamese sub-neural networks. Step S101 includes:
[0079] Step 1, obtain the template image and the search image through the Siamese sub-neural networks.
[0080] Step 2, input the template image into the ResNet50 network of the Siamese sub-neural networks to obtain the vector convolution operation feature, two-dimensional matrix convolution operation feature, three-dimensional matrix convolution operation feature, four-dimensional matrix convolution operation feature, and five-dimensional matrix convolution operation feature of the template image, which are respectively f t1 、f t2 、f t3 、f t4 、f t5, as a template branch feature.
[0081] Step 3: Input the search image into the twin neural network ResNet50 network to obtain the vector convolution operation features, two-dimensional matrix convolution operation features, three-dimensional matrix convolution operation features, four-dimensional matrix convolution operation features, and five-dimensional matrix convolution operation features of the search image, which are f s1 , f s2 , f s3 , f s4 , f s5 , as the search branch features.
[0082] The size of the search image is larger than that of the template image.
[0083] In one embodiment, the twin neural network includes a multi-branch fusion module, and step S102 includes:
[0084] Step 1: Compress the channels of the features f t3 , f t4 , f t5 of the template branch feature to obtain the features f n3 , f n4 , f n5 .
[0085] Step 2: Pass the feature f t2 of the template branch feature through the multi-branch fusion module to obtain the feature f n2 with different receptive fields;
[0086] Step 3: Add f n3 , f n4 , f n5 to f n2 respectively and perform a central cropping operation to obtain the template features F t3 , F t4 , F t5 .
[0087] In one implementation, perform 1×1 convolutions on the features f t3 , f t4 , f t5 of the search branch feature respectively, and compress the number of channels to 256 to obtain the features f n3 , f n4 , f n5 . The sizes of f n3 , f n4 , f n5 are all 15×15×256. The sizes of the template features F t3 , F t4 , F t5 of the template image are 7×7×256.
[0088] See Figure 3 , Figure 3 , which is a block diagram of the multi-branch fusion module provided by the embodiment of the present invention.
[0089] The working steps of the multi-branch fusion module are as follows: First step, input the above two-dimensional matrix convolution operation feature into a two-step convolution kernel (the size of the two-step convolution kernel is 3×3, and the stride is 1). Through two-step convolution operations, a feature map with a size of 31×31×128 (hereinafter referred to as the first feature map) is output. Second step, input the first feature map into two branches. The first branch keeps the input feature unchanged, and the other branch is two identical convolutional sub-networks. Through the above operations, features with different receptive fields of the two-dimensional matrix convolution operation feature can be obtained, and deeper features containing multiple semantic information of the target can be obtained. Third step, connect the features of the two branches. Finally, perform a downsampling operation to obtain a refined feature map f n2 .
[0090] Add f n2 to f n3 , f n4 , f n5 respectively and then perform a central cropping operation, as specifically shown in the following formula (1):
[0091] F t3 = Crop(f n3 + f n2 )
[0092] F t4 = Crop(f n4 + f n2 )(1)
[0093] F t5 = Crop(f n5 + f n2 )
[0094] In one embodiment, the siamese neural network includes a global context module, and step S103 includes:
[0095] Step one, perform channel compression on the features of f s3 , f s4 , f s5 of the search branch feature to obtain features f m3 , f m4 , f m5 .
[0096] Step two, pass f m3 , f m4 , f m5 through the global context module to obtain search features F s3 , F s4 , Fs5 。
[0097] In one implementation, the f of the search branch features s3 、f s4 、f s5 features are respectively convolved with 1×1 to compress the number of channels to 256 each, obtaining the features f m3 、f m4 、f m5 。The sizes of f m3 、f m4 、f m5 are all 15×15×256. The sizes of the search features F s3 、F s4 、F s5 are 31×31×256.
[0098] See Figure 4 , Figure 4 which is the block diagram of the global context module provided by the embodiment of the present invention.
[0099] The global context module includes three parts: a context modeling sub-module, a transformation sub-module, and a fusion sub-module. Assume that x and z respectively represent the input and output of the global context module, N p represents the number of elements in the feature map.
[0100] The global context operation can be represented by formula (2), where W1, W2, and W3 respectively represent Figure 4 the weight coefficients of the three convolutions with kernel sizes in , LN() represents the layer normalization function for normalization, and ReLu() represents the piecewise linear function for unilateral suppression.
[0101]
[0102] In one embodiment, the siamese neural network includes an attention map module, and step S104 is specifically:
[0103] Input the template features F t3 、F t4 、F t5 and the search features F s3 、F s4 、F s5 into the self-attention module and the cross-attention module of the attention map module respectively, obtaining the attention maps of the template features and the attention maps
[0104] SeeFigure 5 , Figure 5 It is a block diagram of the attention map module provided by the embodiment of the present invention.
[0105] As Figure 5 shown, the attention map module includes a self-attention module and a cross-attention module. In order to learn more refined semantic features from space and channels, self-attention and cross-attention sub-networks are proposed. As Figure 5 shown, there are 4 dashed boxes from top to bottom. Among them, the content in the first dashed box and the fourth dashed box represents self-attention, and the content in the second dashed box and the third dashed box represents cross-attention. Specifically, they are denoted as the template feature Z and the search feature X respectively, where the feature sizes of Z and X are C×h×w and C×H×W respectively.
[0106] The self-attention module is composed of spatial attention and channel attention.
[0107] For spatial attention, first, the search feature X is segmented according to spatial positions to obtain where X i,j ∈R C×1×1 corresponding to the parameter of the spatial position (i, j); secondly, 1×1 convolution is used to compress the channels, and the corresponding formula is Q = W sq *X, where W sq ∈R C×1×1×1 is the parameter of the convolution kernel, and Q∈R H×W is obtained. At the same time, the value of each spatial position of Q can be expressed as Then, a feature with spatial information is generated, where σ() is the sigmoid activation function; finally, X~ is given a learnable parameter α and added to the original feature X to obtain the final feature X sa , as shown in the following formula (3):
[0108]
[0109] For channel attention, first, the input feature X is separated according to the number of channels to obtain x i ∈R H ×W ; secondly, global average pooling operation is performed on the space to generate a vector V∈R C×1×1 , and the value of the k-th channel can be obtained by the following formula (4):
[0110]
[0111] Again, two convolution operations are used to compress and expand V to obtain
[0112] Among them, and are the parameters corresponding to two convolutional kernels respectively; then, a feature map with channel aggregation features is obtained where σ() is the sigmoid activation function; finally, is given a learnable parameter β and added to the original feature X to obtain the final channel feature Figure X ca , as shown in the following formula (5):
[0113]
[0114] Cross-attention module
[0115] For the search branch, the template feature Z and the search feature X are input into the cross-attention module, which is located in Figure 5 the second dashed box and the third dashed box. Then the template feature undergoes a global average pooling and two 1×1 convolutional operations, and finally a channel feature map is obtained where C is the number of channels; secondly, is subjected to a sigmoid function activation operation and multiplied by the initial feature x i to obtain the preliminary feature as shown in the following formula (6):
[0116]
[0117] Again, the corresponding features are calculated to calculate the final cross-attention feature map Then, the cross-attention feature Figure X cro is obtained from X and as follows:
[0118]
[0119] λ in the above formula (7) is a learnable parameter.
[0120] The features X sa , X ca and X cro are fused in parallel through element-wise summation operations, and the attention search feature map is effectively obtained. The acquisition of the corresponding attention features in the template feature is the same as that of the search branch.
[0121] In one embodiment, the Siamese neural network includes a depth cross-correlation module, and step S105 is specifically:
[0122] Through the depth cross-correlation module, the attention map of the template feature and the attention maps of the search features respectively perform depth cross - correlation operations to obtain score maps φ3, φ4, and φ5.
[0123] In one implementation, the attention map of the template feature and the attention maps of the search features perform depth cross - correlation operations to obtain φ3, and so on to obtain φ4 and φ5.
[0124] In one embodiment, the Siamese neural network includes a target position determination module, and step S106 includes:
[0125] Step 1, input the score maps φ3, φ4, and φ5 into the classification branch and the regression branch of the target position determination module respectively.
[0126] Step 2, through the classification branch, the score maps φ3, φ4, and φ5 respectively pass through a convolution with a convolution kernel size of 1×1 and a stride of 1 to obtain features with 2K channels Multiply with the preset learnable weights respectively to obtain classification features The classification features include the foreground and background features of the target in the search image.
[0127] Step 3, through the regression branch, the score maps φ3, φ4, and φ5 respectively pass through a convolution with a convolution kernel size of 1×1 and a stride of 1 to obtain features with 4k channels Multiply with the preset learnable weights respectively to obtain regression features The regression features include the features of the target.
[0128] Step 4, determine the position of the target in the search image according to the classification features and the regression features In one implementation,
[0129] can be multiplied with different learnable weight coefficients respectively, that is can be multiplied with different learnable weight coefficients respectively, that is can be can be multiplied with different learnable weight coefficients respectively, that is
[0130] See Figure 6 , Figure 6 which is the accuracy and success rate test chart of the target tracking method provided by the embodiment of the present invention.
[0131] The target tracking method that fuses significant information and multi-granularity context features provided by the embodiments of the present invention is tested with nine other mainstream target tracking methods on the OTB2015 dataset to obtain the accuracy test chart (a) and the success rate test chart (b). It can be seen from Figure 6 that the target tracking method provided by the embodiments of the present invention is optimal in terms of accuracy and success rate.
[0132] See Figure 7 , Figure 7 which is the EAO value test chart of the target tracking method provided by the embodiments of the present invention.
[0133] The EAO value test chart obtained by testing the target tracking method that fuses significant information and multi-granularity context features provided by the embodiments of the present invention with the current mainstream target tracking methods on the VOT2019 dataset. The larger the EAO value, the better the evaluation effect. It can be seen from Figure 7 that the target tracking method of the present invention has the optimal performance among many target tracking methods.
[0134] See Figure 8 , Figure 8 which is the EAO value test chart of the target tracking method provided by the embodiments of the present invention for various situations.
[0135] The target tracking method that fuses significant information and multi-granularity context features provided by the embodiments of the present invention is tested with other mainstream target tracking methods on VOT2019 to obtain the EAO values in various situations, including camera motion, occlusion, size change, illumination change, and motion change. See Figure 8 and it can be seen that the target tracking method of the present invention shows good performance when facing camera motion, illumination change, and motion change.
[0136] See Figure 9 , Figure 9 which is the representative visual result chart of the target tracking method provided by the embodiments of the present invention.
[0137] Figure 9Shows the representative visual results of different object tracking methods on the OTB2015 dataset, that is, 10 representative video sequences are selected from the OTB2015 dataset, and in these video sequences, the object tracking method of the present invention is compared with other mainstream object tracking methods. The mainstream object tracking methods for comparison include: Ocean, MDNet, DaSiamRPN, ATOM, SiamRPN++, and SiamBAN. When facing the situations of motion blur, fast motion, and low-resolution object tracking, the object tracking method of the present invention shows more accurate tracking accuracy. For example, in the three video sequences of BlurOwl, soccer, and DragonBaby, some object tracking methods suffered losses, but the object tracking method of the present invention is more robust in tracking while maintaining a high degree of accuracy. At the same time, it can also be seen from other video sequences in the figure that the object tracking method of the present invention also shows better tracking performance when facing object tracking situations such as rotation, scale change, deformation, and occlusion.
[0138] The embodiment of the present invention also provides an electronic device, such as Figure 10 shown, including a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004.
[0139] The memory 1003 is used to store computer programs;
[0140] The processor 1001, when executing the program stored on the memory 1003, implements the above-mentioned object tracking method that fuses significant information and multi-granularity context features.
[0141] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0142] The communication interface is used for communication between the above electronic device and other devices.
[0143] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0144] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0145] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned object tracking methods that fuse significant information and multi-granularity context features are implemented.
[0146] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the object tracking methods that fuse significant information and multi-granularity context features in the above embodiments.
[0147] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).
[0148] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0149] Each embodiment in this specification is described in a related manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the system, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0150] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.
Claims
1. An object tracking method that fuses significant information and multi-granularity context features, characterized in that, The method is applied to a siamese neural network, and is characterized in that the method comprises: Obtain a template image and a search image, extract multiple features of the template image as template branch features, and extract multiple features of the search image as search branch features; the template image contains the appearance information of the target to be tracked; the search image is an image containing the target; Obtain the template feature of the template image according to the template branch features; Obtain the search feature of the search image according to the search branch features; According to the search feature and the template feature, obtain the attention map of the search feature and the attention map of the template feature; Perform depth cross-correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map; Perform classification and regression operations on the score map to determine the position of the target in the search image; The siamese neural network includes siamese sub-neural networks. The obtaining of the template image and the search image, the extraction of multiple features of the template image as template branch features, and the extraction of multiple features of the search image as search branch features include: Obtain the template image and the search image through the siamese sub-neural networks; the size of the search image is larger than the size of the template image; Input the template image into the ResNet50 network of the twin neural network, and the vector convolution operation feature, two-dimensional matrix convolution operation feature, three-dimensional matrix convolution operation feature, four-dimensional matrix convolution operation feature, and five-dimensional matrix convolution operation feature of the template image are respectively f t1 、f t2 、f t3 、f t4 、f t5 , which are used as the template branch features; Input the search image into the twin neural network ResNet50 network, and the vector convolution operation feature, two-dimensional matrix convolution operation feature, three-dimensional matrix convolution operation feature, four-dimensional matrix convolution operation feature, and five-dimensional matrix convolution operation feature of the search image are respectively f s1 、f s2 、f s3 、f s4 、f s5 , which are used as the search branch features; The siamese neural network includes a multi-branch fusion module. The obtaining of the template feature of the template image according to the template branch features includes: Channel compress the features f t3 , f t4 , f t5 of the template branch feature to obtain the features f n3 , f n4 , f n5 ; The f of the template branch feature t2 The feature passes through the multi-branch fusion module to obtain a feature f with different receptive fields n2 ; Add f n3 , f n4 , and f n5 to f s2 respectively, and perform a central cropping operation to obtain the template features F t3 , F t4 , and F t5 of the template image.
2. The method according to claim 1, wherein The siamese neural network includes a global context module. The obtaining of the search feature of the search image according to the search branch features includes: The f of the search branch feature s3 , f s4 , f s5 features are subjected to channel compression to obtain the feature f m3 , f m4 , f m5 ; The f m3 , f m4 , f m5 features pass through the global context module to obtain the search features F s3 , F s4 , F s5 .
3. The method according to claim 2, wherein The siamese neural network includes an attention map module. The obtaining of the attention map of the search feature and the attention map of the template feature according to the search feature and the template feature includes: Input the template features F t3 、F t4 、F t5 and the search features F s3 、F s4 、F s5 into the self-attention module and cross-attention module of the attention map module respectively, to obtain the attention maps of the template features 、 、 and the attention maps of the search features 、 、 。 4. The method according to claim 3, wherein The siamese neural network includes a depth cross-correlation module. The performing of depth cross-correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map includes: Through the depth cross - correlation module, the attention maps of the template features , , and the attention maps of the search features , , are respectively subjected to depth cross - correlation operations to obtain score maps Φ3, Φ4, and Φ5.
5. The method according to claim 4, characterized in that The siamese neural network includes a target position determination module. The performing of classification and regression operations on the score map to determine the position of the target in the search image includes: Input the score maps Φ3, Φ4, and Φ5 into the classification branch and the regression branch of the target position determination module respectively; The score maps Φ3, Φ4, and Φ5 are respectively convolved through a classification branch with a convolution kernel of size 1×1 and a stride of 1 to obtain features with 2K channels. , , , and the preset learnable weights are respectively multiplied by , , to obtain classification features ; the classification features include the foreground and background features of the target in the search image. The score maps Φ3, Φ4, Φ5 are respectively convolved through a convolution kernel with a size of 1×1 and a stride of 1 via a regression branch to obtain features with 4k channels , , The preset learnable weights are respectively multiplied by , , to obtain regression features ; the regression features contain the features of the target; According to the classification features and the regression features determine the position of the target in the search image.
6. An electronic device, characterized in that, Comprising a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used for storing a computer program; when the processor executes the program stored on the memory, it implements the method steps described in any one of claims 2-5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of claims 2-5.
8. An object tracking system that fuses significant information and multi-granularity context features for implementing the method according to claim 1, characterized in that, Comprising siamese sub-neural networks, a multi-branch fusion module, a global context module, an attention map module, a depth cross-correlation module, and a target position determination module, wherein: The twin neural network is used to obtain a template image and a search image, extract multiple features of the template image as template branch features, and extract multiple features of the search image as search branch features; the template image contains the appearance information of the target to be tracked; the search image is an image containing the target. The multi-branch fusion module is used to obtain the template feature of the template image according to the template branch features. The global context module is used to obtain the search feature of the search image according to the search branch features. The attention map module is used to obtain the attention map of the search feature and the attention map of the template feature according to the search feature and the template feature. The depth cross-correlation module is used to perform depth cross-correlation on the attention map of the template feature and the attention map of the search feature to obtain a score map. The target position determination module is used to perform classification and regression operations on the score map to determine the position of the target in the search image.
Citation Information
Patent Citations
Video action recognition method based on residual 3D CNN and multi-modal feature fusion strategy
CN111325155A
CNN-based saliency detection system and method
CN112927209A