Twin Tracking Method and System Optimized Based on Interactive and Aggregative Features
By building a gated dual-view aggregation module, a semantic-guided attention module and a correlation graph aggregation module, the significant characteristics of the target are enhanced, and the robustness and accuracy problems of the twin tracking algorithm in complex environments is solved, and a more stable target tracking effect is achieved.
Patent Information
- Application Number
- CN202210600748.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The existing twin tracking algorithm lacks effective solutions to scenes such as similar interference, complex backgrounds and occlusion in complex environments, resulting in unstable and insufficient robustness of tracking results.
Using a twin tracking method based on interaction and aggregation feature optimization, the gated dual-view aggregation module, semantic-guided attention module and correlation map aggregation module are constructed to enhance the significant features of the target, suppress background noise, and realize self-attention aggregation and interaction of templates and search area features.
Obtain more stable, robust and accurate tracking results in complex environments, improving the accuracy and anti-interference ability of target tracking.
Smart Images

Figure CN114926652B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a Siamese tracking method and system based on interactive and aggregative feature optimization. Background Art
[0002] In the field of computer vision, object tracking is one of the most important and active research topics, which has extensive applications in aspects such as unmanned driving, intelligent security, human-computer interaction, and unmanned aerial vehicles. For a single object tracker or tracking system, it aims to continuously predict the spatial position of an object in subsequent video sequences based on the given arbitrary object coordinate information in the first frame.
[0003] In recent years, the application of siamese networks in the field of object tracking has made great progress. Benefiting from deep learning, it has good feature expressiveness, and can achieve real-time tracking speed through means such as parameter sharing and offline training, and is becoming the mainstream of research in the tracking field. The basic idea of the tracking algorithm based on siamese networks is as follows: the target area corresponding to the ground truth box in the first frame of the video is used as the template, and the subsequent frames are used as the search area. The tracking process is to match the area in the search area that is most similar to the template and use it as the predicted position of the target in the current frame. SiamFC (Bertinetto L, Valmadre J, Henriques J F, et al. Fully-convolutional siamese networks for object tracking. Proceedings of the European Conference on Computer Vision Workshops. 2016, pp. 850-865.) and SiamRPN (Li B, Yan J, Wu W, et al. High-performance visual tracking with siamese region proposal network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018, pp. 8971-8980.) obtain an appearance model through offline training on a large dataset and do not update parameters during online tracking. Therefore, the tracker not only has high accuracy but also has an advantage in terms of speed. However, due to the fixed template, the tracker is not particularly sensitive to changes in the target appearance and is also easily interfered by similar objects and complex backgrounds.To adapt to the appearance changes of the target, CFNet (Valmadre J, Bertinetto L, Henriques JF, et al. End-to-end representation learning for correlation filter based tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017, pp. 5000-5008.) and RASNet (Wang Q, Teng Z, Xing J, et al. Learning attentions: Residual attentional siamese network for high performance online visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018, pp. 4854-4863.) respectively embed the correlation filter module and the attention mechanism module in the template branch, and update the template parameters through local modules. GradNet (Li P, Chen B, Ouyang W, et al. Gradnet: Gradient-guided network for visual object tracking. Proceedings of the IEEE International Conference on Computer Vision. 2019, pp. 6161-6170.) and UpdateNet (Zhang L, Abel G, Joost V, et al. Learning the model update for siamese trackers. Proceedings of the IEEE International Conference on Computer Vision. 2019, pp. 4009-4018.) use the process of network iterative learning to achieve template parameter update.Compared with directly updating parameters on the first-frame template in GradNet and UpdateNet, MemDTC (Yang T, Antoni B. Visual tracking via dynamic memory networks. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2021, pp. 360-374.) stores reliable target templates during the tracking process by adding memory units. Therefore, it can completely preserve the effective information of the first-frame template, which helps the tracker quickly recover when it drifts. In addition, to improve the discriminative ability of the siamese tracker for similar objects and complex backgrounds, DaSiamRPN (Zhu Z, Wang Q, Li B, et al. Distractor-aware siamese networks for visual object tracking. Proceedings of the European Conference on Computer Vision. 2018, pp. 103-119.) online designs a distractor-aware module that can perform incremental learning. Nocal-Siam (Tan H, Zhang X, Zhang Z, et al. Nocal-siam: Refining visual features and response with advanced non-local blocks for real-time siamese tracking. IEEE Transactions on Image Processing. 2021, pp. 30:2656-2668.) utilizes the long-distance dependence of non-local attention to strengthen the learning of feature weights related to the target. SiamDW (Zhang Z, Peng H. Deeper and wider siamese networks for real-time visual tracking. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2019, pp. 4591-4600.) designs a deeper and wider network architecture for the siamese tracker to further explore the feature extraction and discriminative capabilities of deep networks.
[0004] Although certain progress has been made in current twin tracking algorithms in aspects such as designing deeper and wider backbone networks, better matching methods, more accurate output representations, and more efficient online update mechanisms, there is still a lack of more effective solutions for scenarios such as similar object interference, complex backgrounds, and occlusions. Summary of the Invention
[0005] The purpose of the present invention is to provide a twin tracking method and system based on interactive and aggregative feature optimization, which is conducive to obtaining more stable, robust, and accurate tracking results in complex environments.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is: a twin tracking method based on interactive and aggregative feature optimization, including the following steps:
[0007] S1. Initialize the template image and the search area image;
[0008] S2. Construct a feature extraction network, input the template image and the search area image, and obtain the corresponding template multi-layer features F z and the search area multi-layer features F x ;
[0009] S3. Construct a gated dual-view aggregation module GDA to optimize the template multi-layer features, input the template multi-layer features F z into the GDA module, and obtain the optimized template multi-layer features
[0010] S4. Construct a semantic-guided attention module SGA to achieve coarse-grained feature optimization of the search area, input and the search area multi-layer features F x into the SGA module, and obtain the coarsely optimized search area features
[0011] S5. Construct a correlation graph aggregation module CGA to achieve fine-grained feature optimization of the search area, input and into the CGA module, and obtain the finely optimized search area features
[0012] S6. Construct a head prediction network, input and to predict the position of the target in the current frame.
[0013] Furthermore, the specific implementation method of step S1 is:
[0014] According to the given target ground truth bounding box in the first frame, a template image with a size of 3×127×127 is cropped from the first-frame image; starting from the second frame, taking the center coordinates of the target predicted bounding box in the previous frame as the reference point, a search region image with a size of 3×255×255 is cropped.
[0015] Furthermore, the specific implementation method of step S2 is as follows:
[0016] Using ResNet-50 as the feature extraction network, taking the template image and the search region image as inputs to obtain the multi-layer features of the template and the multi-layer features of the search region where l represents the total number of layers of the template or search region features extracted, represent the template feature and the search region feature of the i-th layer respectively, i ∈ [1, l].
[0017] Furthermore, the specific implementation method of step S3 is as follows:
[0018] The GDA module includes three sub-modules: local perspective attention LA, global perspective attention GA, and aggregation gating; the LA module is used to highlight the high-frequency information of the local perspective; for a single-layer template feature of size C×H×W The local perspective attention feature is expressed as:
[0019]
[0020] where W2 is a learnable convolution parameter with a size of and here r represents the channel compression parameter; represents batch normalization; σ represents the sigmoid function; represents element-wise multiplication; the high-frequency feature is obtained by subtracting the local mean, expressed as:
[0021]
[0022]
[0023] In the formula, W1 is a learnable convolution parameter; is the feature mapped by the convolution of W1; AvgPool(·) represents average pooling, used to obtain the average signal intensity of the local region; ks and s represent the window size and the stride respectively; δ represents the non-linear activation function, here ReLU is adopted;
[0024] The LA module focuses on a fixed receptive field and aggregates information in a local area through convolutional operations; while the GA module is used to aggregate global information with different receptive fields through the interaction of multi-layer features. For a set of l-layer features F = {x1, x2,..., x l}, for any two layers of features and First, three convolutional layers θ(·), φ(·), and g(·) are used to perform linear mapping on x i to obtain feature maps of "query", "key", and "value" and That is
[0025] Q = θ(x i )
[0026] K1 = φ(x i )
[0027] V1 = g(x i )
[0028] At the same time, the feature x j shares the convolutional layers φ(·) and g(·) to obtain the corresponding feature maps and That is
[0029] K2 = φ(x j )
[0030] V2 = g(x j )
[0031] Then, the "keys" and "values" of each layer are concatenated together respectively to obtain the global representation of multi-layer features and where S = l × H × W, here l represents the total number of feature layers to be queried; thus the global features K and V are expressed as:
[0032] K = [φ(x i ) || φ(x j )]
[0033] V = [g(x i ) || g(x j )]
[0034] where, [· || ·] represents concatenation of features in the spatial dimension;
[0035] Therefore, according to the formula of the standard non-local attention, the attention feature y i is obtained and expressed as:
[0036]
[0037] In the above formula, it represents performing a softmax operation along the j dimension;
[0038] Finally, the channels of y are unified with the original feature map through a convolutional layer ξ(·), and added to the original feature map in a residual manner to obtain: i and the channels of the original feature map, and added to the original feature map in a residual manner to obtain:
[0039]
[0040] For the template multi-layer features According to the above formula, the i-th layer feature After being updated by the GA module, it is expressed as:
[0041]
[0042] Among them,
[0043]
[0044] The LA module and the GA module reduce the channel redundancy to a certain extent through the channel compression parameters r and m, and obtain the attention features from the local and global perspectives respectively and On this basis, the gating mechanism of the aggregation gating module is used to adaptively fuse and Thereby enhancing the effective representation of the significant features; for the input features and First, the two features are concatenated together, and a 1×1 convolutional layer is used to learn the correlation between them, and then the sigmoid function is used to obtain the normalized correlation matrix W gate , let it represent the weight matrix of, then the weight matrix of is expressed as 1 - W gate ; then, feature weighting is achieved by multiplying the weight matrix with the features element by element, and the finally obtained optimized feature is expressed as:
[0045]
[0046] Furthermore, the specific implementation method of step S4 is:
[0047] For the optimized feature output by the GDA module The SGA module extracts the global semantic information in the spatial dimension layer by layer, generates the target semantic attention matrix, and then interacts with the search area feature layer by layer to obtain the coarsely optimized search area feature Specifically, for the i-th layer template feature The generated target semantic attention matrix is expressed as:
[0048]
[0049] where GAP(·) represents global average pooling in the spatial dimension, and σ is the sigmoid function;
[0050] Then, the SGA module aggregates multi-layer features with a global perspective attention, and it shares parameters with the global perspective attention in the GDA module to reduce the actual computational amount; through the interaction of the target semantic information with the "query", "key", and "value" features of the search region, the feature of the i-th layer search region The generated Q, K, and V are respectively expressed as:
[0051]
[0052]
[0053]
[0054] So the optimized feature obtained by this layer is expressed as:
[0055]
[0056] where
[0057]
[0058] In the above formula, i represents the current layer, and j represents other multi-layer features.
[0059] Furthermore, the specific implementation method of step S5 is:
[0060] The CGA module takes and as inputs. On the one hand, it calculates the correlation between the spatial pixels of the search region and the template globally; on the other hand, it calculates the local correlation only based on the salient features of the template; by fusing the global and local correlations, it uses graph convolution to strengthen the connection of spatial positions, so as to realize the construction of the attention map associated with the target features; specifically, for the optimized feature of the template of the i-th layer and the optimized feature of the search region First, the template is sliced along the space and channels respectively to obtain the spatial feature and the channel feature where N1 = H1×W1; for a certain pixel in the search region, first calculate its correlation with the spatial feature of the template to obtain the spatial correlation map S1, which is expressed as:
[0061]
[0062] Among them, Corr(·) is a correlation calculation function, and the inner product method is adopted here;
[0063] Next, based on S1, the global information of the template is retrieved by calculating the correlation with the channel features. At this time, the calculated correlation map S2 is expressed as:
[0064]
[0065] The correlation between a certain pixel in the search area and the template global is expressed as:
[0066]
[0067]
[0068] Among them, MaxPool(·) is the max pooling operation, and ks and s represent the window size and the stride respectively;
[0069] So the local correlation is expressed as:
[0070]
[0071] Finally, the two correlation maps are fused by the method of adding corresponding elements, and their graph relationship is constructed to enhance the position association; the obtained correlation map and the search area feature are added to obtain a finer-grained optimized feature which is expressed as:
[0072]
[0073] Among them, GCN(·) is a two-layer graph convolutional network; denotes adding corresponding elements; and then a finer-grained optimized search area feature is obtained
[0074] The present invention also provides a Siamese tracking system based on interactive and aggregative feature optimization, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above method steps can be implemented.
[0075] Compared with the prior art, the present invention has the following beneficial effects: The method and system enhance the target salient features and suppress background noise through self-attention aggregation and interaction of template features and search region features. Specifically, a novel interaction and aggregation network is adopted, which includes a gated dual-view attention module, a semantic-guided attention module, and a correlation graph aggregation module. Among them, the gated dual-view attention module aggregates the outputs of two sub-modules of local-view attention and global-view attention based on a gating mechanism, and is used to enhance the features of target saliency and discriminability. The semantic-guided attention module extracts the semantic information of the target and uses it as a priori to guide the feature optimization of the search region. Further, for the optimized template and search region features, local and global similarities are respectively constructed in the correlation graph aggregation module, and the spatial position relationship is strengthened through a graph convolutional network. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is the flowchart of the method implementation of the embodiment of the present invention.
[0077] Figure 2 is the precision comparison result of the method of the present invention and other object tracking methods under different attributes in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0079] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0080] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0081] As Figure 1 shown, this embodiment provides a Siamese tracking method based on interactive and aggregative feature optimization, including the following steps:
[0082] S1. Initialize the template image and the search region image. The specific implementation method is as follows:
[0083] According to the given ground-truth bounding box in the first frame, a template image with a size of 3×127×127 is cropped from the first-frame image; starting from the second frame, using the center coordinates of the target predicted bounding box in the previous frame as the reference point, a search region image with a size of 3×255×255 is cropped. The target predicted bounding box refers to the predicted target position in each frame, given in the form of (x, y, w, h), where (x, y) represents the center position of the predicted bounding box, and w and h represent the width and height of the predicted bounding box respectively.
[0084] S2. Construct a feature extraction network, input the template image and the search region image, and obtain the corresponding multi-layer template features and multi-layer search region features The specific implementation method is as follows:
[0085] Use ResNet-50 and its improved network as the feature extraction network, take the template image and the search region image as inputs, and obtain the multi-layer template features and multi-layer search region features where l represents the total number of layers of the template or search region features extracted, represent the template feature and the search region feature of the i-th layer (i ∈ [1, l]) respectively.
[0086] S3. Construct a gated dual-view aggregation module GDA to optimize the multi-layer template features, input the multi-layer template features into the GDA module, and obtain the optimized multi-layer template features The specific implementation method is as follows:
[0087] 1) The GDA module includes three sub-modules: local-view attention (LA), global-view attention (GA), and aggregation gate. The LA module is used to highlight the high-frequency information of the local view. To save computational parameters, a bottleneck structure scheme is adopted for design, and the detailed features are highlighted through residual connections to protect the original features and prevent the effective information from being wrongly erased. In addition, for the attribute of convolutional local connection, essentially each element of the feature map represents the embedded feature and signal strength of a specific region in the previous layer feature map. Therefore, high-frequency discriminative information can be obtained by subtracting the average signal strength. When the difference between the target and the background is not obvious, enhancing the discriminative high-frequency information representation is crucial. Thus, for a single-layer template feature of size C×H×W The local-view attention feature is expressed as:
[0088]
[0089] Among them, W2 is a learnable convolution parameter, with a size of and where r represents the channel compression parameter; represents batch normalization; σ represents the sigmoid function; represents element-wise multiplication; high-frequency features are obtained by subtracting the local mean, denoted as:
[0090]
[0091]
[0092] In the formula, W1 is a learnable convolution parameter; is the feature mapped by W1 convolution; AvgPool(·) represents average pooling, which is used to obtain the average signal intensity of the local area; ks and s represent the window size and stride respectively; δ represents the non-linear activation function, and ReLU is adopted here.
[0093] 2) The LA module focuses on a fixed receptive field and aggregates the information of the local area through convolution operations; while the GA module is used to aggregate the global information of different receptive fields through the interaction of multi-layer features. The standard non-local attention captures the dependencies at different positions through pixel-level interaction, greatly improving the performance of visual tasks. However, it does not consider the information interaction between different network layers and ignores the importance of different receptive fields for semantic information mining. For the above considerations, the standard non-local attention is extended to cross-layer non-local attention, which aggregates multi-layer semantic information to the current layer through interaction with different receptive fields, so as to obtain a richer feature representation. For the sake of convenience of expression and general generalization, for a set of l-layer features F = {x1, x2,..., x l}, for any two-layer features and First, three convolutional layers θ(·), φ(·) and g(·) are used to linearly map x i to obtain the feature maps of "query", "key" and "value" and That is
[0094] Q = θ(x i )
[0095] K1 = φ(x i )
[0096] V1 = g(x i )
[0097] Meanwhile, feature x j Share the convolutional layers φ(·) and g(·) to obtain the corresponding feature maps and That is
[0098] K2 = φ(x j )
[0099] V2 = g(x j )
[0100] Then, concatenate the "keys" and "values" of each layer together to obtain the global representation of the multi-layer features and where S = l × H × W, and here l represents the total number of feature layers to be queried. Thus, the global features K and V are expressed as:
[0101] K = [φ(x i ) || φ(x j )]
[0102] V = [g(x i ) || g(x j )]
[0103] where [· || ·] represents concatenation of features along the spatial dimension.
[0104] Therefore, according to the formula of the standard non-local attention, the attention feature y i is expressed as:
[0105]
[0106] In the above formula, it represents the softmax operation along the j dimension.
[0107] Finally, pass the feature y i through a convolutional layer ξ(·) to unify the number of channels with the original feature map and add it to the original feature map in a residual manner, thus obtaining:
[0108]
[0109] For the template multi-layer features According to the above formula, the updated representation of the i-th layer feature after passing through the GA module is:
[0110]
[0111] where,
[0112]
[0113] 3) The LA module and the GA module reduce the channel redundancy to a certain extent through the channel compression parameters r and m, and obtain the attention features from local and global perspectives respectively. and On this basis, the gating mechanism of the aggregation gating module is used to adaptively fuse and Thereby enhancing the effective representation of significant features. For the input features and First, the two features are concatenated together, and a 1×1 convolutional layer is used to learn the correlation between them, and then the sigmoid function is used to obtain the normalized correlation matrix W gate , let it represent The weight matrix of The weight matrix of is represented as 1 - W gate . Then, feature weighting is achieved by element-wise multiplication (Hardamard product) of the weight matrix and the features, and the finally obtained optimized feature is represented as:
[0114]
[0115] S4. Construct a semantic-guided attention module SGA to achieve coarse-grained feature optimization of the search area. Input and the multi-layer features of the search area into the SGA module to obtain the coarsely optimized search area features The specific implementation method is as follows:
[0116] 1) For the optimized features output by the GDA module, the SGA module extracts the global semantic information in the spatial dimension layer by layer, generates the target semantic attention matrix, and then interacts with the search area features layer by layer to obtain the coarsely optimized search area features Specifically, for the target semantic attention matrix generated by the i-th layer template feature is represented as:
[0117]
[0118] Among them, GAP(·) represents global average pooling in the spatial dimension, and σ is the sigmoid function.
[0119] 2) The SGA module aggregates multi-layer features with a global perspective attention. It shares parameters with the global perspective attention in the GDA module to reduce the actual computational amount; specifically, after the interaction of the target semantic information with the "query", "key", and "value" features of the search area, the i-th layer search area feature The generated Q, K, and V are respectively represented as:
[0120]
[0121]
[0122]
[0123] Then the optimized features obtained by this layer are represented as:
[0124]
[0125] Among them,
[0126]
[0127] In the above formula, i represents the current layer, and j represents the features of other multiple layers.
[0128] S5. Construct a correlation graph aggregation module CGA to achieve fine-grained feature optimization of the search area. Input and into the CGA module to obtain the fine-grained optimized search area features The specific implementation method is as follows:
[0129] The CGA module takes and as inputs. On the one hand, calculate the correlation between the spatial pixels of the search area and the global template; on the other hand, calculate the local correlation only based on the salient features of the template; by fusing the global and local correlations, use graph convolution to strengthen the connection of spatial positions, so as to realize the construction of the attention map associated with the target features.
[0130] 1) For the template optimized features of the i-th layer and the search area optimized features First, split the template along the space and channels respectively to obtain the spatial feature and the channel feature where N1 = H1 × W1; for a certain pixel in the search area, first calculate its correlation with the template spatial feature to obtain the spatial correlation graph S1, which is represented as:
[0131]
[0132] Among them, Corr(·) is the correlation calculation function, and here the inner product method is adopted.
[0133] Then, based on S1, retrieve the global information of the template by calculating the correlation with the channel feature. At this time, the calculated correlation graph S2 is represented as:
[0134]
[0135] For the sake of simplicity of description, the correlation between a certain pixel in the search area and the entire template is expressed as:
[0136]
[0137]
[0138] where MaxPool(·) is the max pooling operation, and ks and s represent the window size and the stride respectively.
[0139] Then the local correlation is expressed as:
[0140]
[0141] 2) The two correlation maps are fused by the method of adding corresponding elements, and the graph relationship is constructed for them, so as to enhance the position association; the obtained correlation map and the search area feature are added to obtain a more fine-grained optimized feature which is expressed as:
[0142]
[0143] where GCN(·) is a two-layer graph convolutional network; represents adding corresponding elements; and then a more fine-grained optimized search area feature is obtained
[0144] S6. Construct a head prediction network, and input and to predict the position of the target in the current frame.
[0145] This embodiment also provides a Siamese tracking system based on interactive and aggregative feature optimization, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the above method steps can be implemented.
[0146] In this embodiment, the OTB100 dataset is used for comparative verification, Figure 2 showing the accuracy comparison results of the method FRIA-Track of the present invention with other object tracking methods under different attributes. Table 1 shows the success rate comparison results of the method proposed in the present invention with other object tracking methods on the OTB100 dataset.
[0147] Table 1 Comparison results between the present invention and other object tracking methods
[0148]
[0149] From Figure 2 It can be seen that the method FRIA-Track of the present invention shows the best level under 8 attributes, and its performance under 10 attributes exceeds that of the baseline algorithm SiamCAR. It can be seen from Table 1 that the method of the present invention has the best success rate compared with other object tracking methods.
[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0154] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in any other form. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A twin tracking method based on interactive and aggregative feature optimization, characterized in that It includes the following steps: S1. Initialize the template image and the search area image; S2. Construct a feature extraction network, input the template image and the search area image, and obtain the corresponding multi-layer features F of the template z and the multi-layer features F of the search area x ; S3. Construct a gated dual-view aggregation module GDA to optimize the multi-level features of the template, and input the multi-level features F of the template into the GDA module to obtain optimized multi-level features of the template z S4. Construct a semantic-guided attention module SGA to achieve fine-grained feature optimization of the search region, and and multi-level features F of the search region x are input into the SGA module to obtain the fine-grained optimized search region features S5. Construct a Correlation Graph Aggregation module CGA to achieve fine-grained feature optimization of the search region. The and are input into the CGA module to obtain the fine-grained optimized search region features S6. Build a head prediction network, and input and to predict the position of the target in the current frame.
2. The twin tracking method based on interactive and aggregative feature optimization according to claim 1, wherein The specific implementation method of step S1 is: According to the target true bounding box given in the first frame, crop a template image with a size of 3×127×127 on the first frame image; starting from the second frame, using the center coordinates of the target predicted bounding box in the previous frame as the reference point, crop a search area image with a size of 3×255×255.
3. The twin tracking method based on interactive and aggregative feature optimization according to claim 1, wherein The specific implementation method of step S2 is: Use ResNet-50 as the feature extraction network, take the template image and the search region image as inputs, and obtain multi-layer features of the template and multi-layer features of the search region where l represents the total number of layers of the template or search region features extracted represent the template feature and the search region feature of the i-th layer respectively, and i ∈ [1, l].
4. The twin tracking method based on interactive and aggregative feature optimization according to claim 1, wherein The specific implementation method of step S3 is: The GDA module includes three sub-modules: Local View Attention (LA), Global View Attention (GA), and Aggregation Gating; the LA module is used to highlight the high-frequency information of the local view; for a single-layer template feature of size C×H×W Local view attention feature It is expressed as: where W2 is a learnable convolution parameter with a size of and where r represents the channel compression parameter; represents batch normalization; σ represents the sigmoid function; represents element-wise multiplication; the high-frequency feature is obtained by subtracting the local mean and is expressed as: Wherein, W1 is a learnable convolution parameter; is the feature mapped by convolution of W1; AvgPool(·) represents average pooling, which is used to obtain the average signal intensity of a local area; ks and s respectively represent the window size and the stride; δ represents a non-linear activation function, and ReLU is adopted here; The LA module focuses on a fixed receptive field and aggregates information in the local area through convolutional operations; while the GA module is used to aggregate global information with different receptive fields through the interaction of multi-layer features. For a set of l-layer features F = {x1, x2,..., x l}, for any two layers of features and First, three convolutional layers θ(·), φ(·), and g(·) are used to perform linear mapping on x i to obtain the feature maps of "query", "key", and "value" and That is Q = θ(x i ) K1 = φ(x i ) V1 = g(x i ) Meanwhile, feature x j Share the convolutional layers φ(·) and g(·) to obtain the corresponding feature maps and That is K2 = φ(x j ) V2 = g(x j ) Then, the "keys" and "values" of each layer are concatenated together respectively to obtain the global representation of the multi-layer features and where S = l×H×W, where l represents the total number of feature layers to be queried; thus, the global features K and V are expressed as: K = [φ(x i ) || φ(x j )] V = [g(x i ) || g(x j )] Among them, [·||·] represents the feature concatenation in the spatial dimension; Therefore, according to the formula of standard non-local attention, the attention feature y is obtained i , which is expressed as: In the above formula, which represents performing a softmax operation along the j dimension; Finally, the number of channels of y is unified with the original feature map through a convolutional layer ξ(·), and added to the original feature map in the form of a residual, resulting in: i For the template multi-layer features According to the above formula, where the feature of the i-th layer The updated representation after passing through the GA module is: Among them, The LA module and the GA module reduce the channel redundancy to a certain extent through the channel compression parameters r and m, and obtain the attention features from local and global perspectives respectively. and On this basis, the gating mechanism of the aggregation gating module is used to adaptively fuse and so as to enhance the effective representation of the salient features; for the input features and First, the two features are concatenated together, and a 1×1 convolutional layer is used to learn the correlation between them, and then the sigmoid function is used to obtain the normalized correlation matrix W gate , let it represent 's weight matrix, then 's weight matrix is represented as 1 - W gate ; then, the feature weighting is realized by multiplying the weight matrix with the features element by element, and the finally obtained optimized feature is expressed as:
5. The twin tracking method based on interactive and aggregative feature optimization according to claim 1, wherein The specific implementation method of step S4 is: For the optimized features output by the GDA module The SGA module extracts the global semantic information in the spatial dimension layer by layer, generates the target semantic attention matrix, and then interacts with the search region features layer by layer to obtain the coarsely optimized search region features Specifically, for the template features of the i-th layer The generated target semantic attention matrix is expressed as: Among them, GAP(·) represents global average pooling in the spatial dimension, and σ is the sigmoid function; Then, the SGA module aggregates multi-layer features using global perspective attention, which shares parameters with the global perspective attention in the GDA module to reduce the actual computational amount; after the interaction of the target semantic information with the "query", "key", and "value" features of the search region, the feature of the search region at the i-th layer The generated Q, K, and V are respectively expressed as: The optimized features obtained for this layer are expressed as: Among them, In the above formula, i represents the current layer, and j represents the features of other multiple layers.
6. The twin tracking method based on interactive and aggregative feature optimization according to claim 1, wherein The specific implementation method of step S5 is: The CGA module takes and as inputs. On the one hand, it calculates the correlation between the spatial pixels in the search area and the global template; on the other hand, it calculates the local correlation only based on the salient features of the template; by fusing the global and local correlations and using graph convolution to strengthen the connection of spatial positions, an attention map associated with the target feature is constructed. Specifically, for the optimized template feature at the i-th layer and the optimized feature of the search area, first, the template is sliced along the space and channels respectively to obtain the spatial feature and the channel feature where N1 = H1 × W1. For a certain pixel in the search area, first calculate its correlation with the spatial feature of the template, and the spatial correlation map S1 is expressed as: Among them, Corr(·) is the correlation calculation function, and here the inner product method is adopted; Then, based on S1, retrieve the global information of the template by calculating the correlation with the channel features. At this time, the calculated correlation map S2 is expressed as: The correlation between a certain pixel in the search area and the template global is expressed as: Among them, MaxPool(·) is the max pooling operation, and ks and s represent the window size and the stride respectively; So the local correlation is expressed as: Finally, the two correlation graphs are fused by adding corresponding elements, and their graph relationships are constructed to enhance the position correlation; the obtained correlation graph is added to the search region feature to obtain a more fine-grained optimized feature which is expressed as: where GCN(·) is a two-layer graph convolutional network; denotes element-wise addition; thus obtaining the fine-grained optimized search region feature 7. A twin tracking system based on interactive and aggregative feature optimization, characterized in that It includes a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor runs the computer program instructions, it can implement the method steps described in any one of claims 1-6.
Citation Information
Patent Citations
Target tracking method and system based on corner attention twin network
CN113297959A
Step-by-step different-source image template matching method based on cascade network
CN114140700A