A Single-Object Tracking Method Based on Siamese Network with Hierarchical Feature Fusion for Anti-Background Interference
Through the hierarchical feature fusion and attention mechanism, combined with the SIOU loss function, the problem of poor target tracking effect in complex backgrounds is solved, and higher tracking accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202410046446.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-01-12
AI Technical Summary
Existing target tracking methods are difficult to distinguish between targets and backgrounds in complex contexts, resulting in poor tracking results.
The anti-background interference twin network method of hierarchical feature fusion is adopted. Low-level detailed information and high-level semantic information are fused through the dual-feature fusion module, and noise is eliminated through the attention focus module, combining with the new SIOU loss function optimization model.
In complex backgrounds, the accuracy and robustness of target tracking are improved, the impact of background interference is reduced, and the model's ability to identify targets is enhanced.
Smart Images

Figure CN118134963B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a single-object tracking method for a Siamese network with hierarchical feature fusion and anti-background interference. Background Art
[0002] Object tracking, as a fundamental task in the field of image processing, compares the target graph provided in the first frame of a video sequence with the similarity of the object in the region of interest in subsequent frames to achieve the tracking of any target, including pedestrians, vehicles, and flying objects. This makes it widely used in autonomous driving, human-computer interaction, and military radar fields. Unfortunately, when encountering complex scenes, the tracking effects of most trackers will be greatly affected. These complex factors include interfering objects similar to the target appearance and backgrounds similar to the target texture. In the object tracking task, there may be other objects similar to the target in appearance, color, or texture features, which may cause the tracking algorithm to misidentify these interfering objects as the target. Therefore, achieving accurate tracking in complex backgrounds is still considered a challenging research goal.
[0003] Benefiting from the fast Fourier transform, the tracking algorithm based on correlation filtering can complete the correlation calculation in a relatively short time, so it has become a classic algorithm in object tracking. In recent years, many research works have been dedicated to improving the tracking algorithm based on correlation filtering to address challenges such as complex backgrounds. Such as methods like ASRCF, CACF, and GFSDCF. These tracking methods based on correlation filters use online update strategies to cope with the appearance changes of the target. However, if the appearance change of the target is similar to the background and there is no effective mechanism to distinguish the target from the background, the tracker may incorrectly update the target template. Subsequently, the correlation filter in subsequent frames may regard the similar background as the target, thus interfering with the tracking result. Therefore, the tracking method based on correlation filtering is not suitable for tracking situations with complex background interference.
[0004] Compared with correlation filter - based trackers, deep - learning - based tracking models extract features at more levels through deep learning, not limited to pixel - level similarity calculation. This endows the Siamese network with good anti - interference characteristics, capable of reducing the impact of background interference on the tracking algorithm. However, while deep networks can extract more discriminative features, they inevitably increase the consumption of computing resources and storage space, and may even lead to difficulties in training and network degradation due to unstable gradients. At the same time, most tracking models (such as SiamFC, SiamRPN) rely only on the features of the last layer extracted by the network for similarity learning. This method itself is a simple linear matching process, resulting in insufficient feature information. The similarity map obtained thereby is vulnerable to interference from similar backgrounds. Moreover, although the image features extracted from the last layer have strong semantic information, they lack edge information and some detailed information, leading to easy tracking failure of the model in the presence of similar interfering objects.
[0005] After retrieval, CN113963032A, a target tracking method based on a Siamese network structure integrating object re - identification, trains a target tracking network model. The target tracking network model includes a classification and regression branch and an object re - identification branch. The classification and regression branch includes a fully - convolutional Siamese network module and a classification and regression module. The backbone network of the fully - convolutional Siamese network module is the same as that of the object re - identification module. The backbone network of the fully - convolutional Siamese network module shares parameters and weights with the backbone network of the object re - identification module. For the video sequence to be tracked, the first frame of the video sequence is used as the template frame, and the tracking target is framed in the first frame. The template frame and each frame after the first frame are respectively used as an image pair and input into the trained target tracking network model to determine the location of the tracking target in the video frame, thus realizing target tracking. The present invention can better improve the ability to distinguish interference from similar targets.
[0006] Patent: A Twin Network Structure Object Tracking Method Incorporating Target Re-identification. The backbone network of the fully convolutional twin network module shares parameters and weights with the backbone network of the target re-identification module. Although this reduces the size and complexity of the model and improves computational efficiency, sharing parameters and weights will make the two modules more coupled, which means that a change in the performance of one module may affect the other module, thus increasing the difficulty of tuning and debugging. The target re-identification module is mainly used to assist in data association, that is, to associate the target's trajectory and event information. Based on this, the present invention creatively proposes an attention focus module. This module realizes feature association through the learning of global features. And by assigning different weight values to the features, the model can focus more on the features related to the target while reducing the interference of unimportant features. This design cleverly avoids a large amount of additional computational complexity and is independent of the network used for classification and regression, reducing the coupling between modules and enabling the model to focus on learning the differences between the target and the background. Furthermore, the input feature map of the attention module is derived from the comprehensive features of the dual feature fusion module. This feature map integrates rich detail information and target motion information, enabling more accurate parsing of the target's motion pattern and behavior. Summary of the Invention
[0007] The present invention aims to solve the problems of the above prior art. A single-object tracking method for a twin network with hierarchical feature fusion and anti-background interference is proposed. The technical solution of the present invention is as follows:
[0008] A single-object tracking method for a twin network with hierarchical feature fusion and anti-background interference, comprising the following steps:
[0009] Step 1: Generate training data: Scale multiple data sets to obtain a template image with a pixel size of 127×127 and a search image with a pixel size of 255×255;
[0010] Step 2: Generate image features: The generated template image and search image pass through the feature extraction network AlexNet. The second-layer low-level features and the last-layer high-level features output by the network are fused through a dual feature fusion module to obtain multi-scale features, and noise introduced by the fusion of features at different layers is eliminated through the attention focus module;
[0011] Step 3: Generate training labels: After performing cross-correlation operation on the search features and the template features, a 17*17 feature response map is obtained. The classification branch determines the label value according to the IOU threshold, and the regression branch generates the label value according to the anchor box coordinate offset. The model learns the position, scale, and category information of the target through the generated labels and makes predictions and inferences based on this information;
[0012] Step 4, Loss function optimization: Input the preprocessed template features and search features into the regression branch and the search branch; use the template features as convolution kernels to perform cross-correlation operations on the search area to obtain their respective feature maps. Calculate the loss of the foreground and background scores on the classification branch similarity map and the classification label of the target using the binary cross-entropy function. The regression branch calculates the loss using the regression offset and the regression label of the target through the new SIOU function; finally, minimize the loss through the stochastic gradient descent algorithm to optimize the model. After several rounds of iteration, select the best tracking model from them.
[0013] Further, the specific steps for generating training data in step 1 are as follows:
[0014] A1. Obtain image pairs containing object annotation information from the training dataset. The image pairs obtained from one sequence are positive sample pairs, which are used to help the model learn the appearance representation and position changes of the target; the image pairs obtained from different sequences are negative sample pairs, which are used for the model to learn how to distinguish the target from the background or other objects.
[0015] A2. Perform central scaling or padding operations on the images, and the formula is as follows:
[0016] s(w + 2p)×s(h + 2p) = A#(1)
[0017]
[0018] Where A represents the size of the scaled image, which is 127 when generating the template image and 255 when generating the search template image, w represents the image width, h represents the image height, s represents the scaling coefficient, and p represents the padding area.
[0019] Further, the specific process for generating image features in step 2 is as follows:
[0020] B1. Input the preprocessed template image and search image into the feature extraction network AlexNet. After five layers of convolution, the sizes of the template feature map and the search feature map are 6*6*256 and 22*22*256 respectively, where 6 and 22 represent the feature map size, and 256 represents the number of channels.
[0021] B2. In order to make the output of the last layer of the feature extraction network the same size as the output of the second layer, perform deconvolution operations on the last layer feature map. The last layer feature F5 output by the feature network will pass through the deconvolution layer, and the calculation formula of the deconvolution layer is as follows:
[0022] F′5 = s(F5 - 1) + 2p - k + 2#(5)
[0023] s = 1 represents the stride, p = 3 represents the padding margin, k = 1 represents the convolutional kernel size, F′5 represents the output feature map size after passing through the convolutional layer, and F5 represents the feature map size input to the convolutional layer. Thus, the 6*6*256 feature map becomes 12*12*256 in size;
[0024] B3. Then, add the F′5 feature map obtained in step C2 to the feature map F2 output by the second layer of the backbone network pixel by pixel to obtain F d , and the specific operation is as follows:
[0025]
[0026] Low-level features provide rich detailed information and background discrimination ability, while high-level features provide semantic information of the target and shape change features.
[0027] Furthermore, the attention focus module in step 2 includes the following steps:
[0028] C1. Contextual attention is used to calculate the similarity between features and generate weights to adjust the importance of features; input the feature F obtained in step B3 d into the attention focus module to obtain F d′ feature, and the specific formula is as follows:
[0029]
[0030] where represents the pixel-wise multiplication operation, σ represents the sigmoid function, and F RCCA () represents the cyclic attention operation. Through this attention focus module, each pixel can better represent the target information, making the tracker more focused on the significant features of the target;
[0031] C2. After obtaining the feature in step C1, pass it through a convolutional layer to reduce the dimension of the input data. The convolutional operation formula is as follows:
[0032]
[0033] where d w and d h represent the width and height of the input data matrix, k w and k h represent the width and height of the convolutional kernel, s represents the moving stride, and p represents the input data padding value; the convolutional operation can model the spatial relationship in the image through the sliding window method of the convolutional kernel; it is used for the model to capture the context information of the area around the target, thereby distinguishing the target from the background.
[0034] Furthermore, in step 3, the generation of training labels is as follows:
[0035] D1. Input the feature dimensions obtained in step C2 into the classification head and the regression head respectively. After a convolution operation, we get 4*4*256 and 20*20*256 respectively. This operation transforms the template features into a low-dimensional representation suitable for the classifier.
[0036] D2. The template images input into the classification branch and the regression branch are upsampled through a convolutional kernel, from 4*4*256 to 4*4*(2k*256) and 4*4*(4k*256). This upsampling operation is used to distinguish the feature patterns of different targets.
[0037] The search features are not upsampled.
[0038] D3. Use the template image features as the convolutional kernel to perform a cross-correlation operation with the search image features. The similarity map obtained by the classification branch is 17*17*2k, and the similarity map obtained by the regression branch is 17*17*4k, where k refers to the number of anchor boxes.
[0039] D4. Generate five anchor boxes with different aspect ratios for each pixel of the 17*17 feature map. The scale distribution is {3, 2, 1, 0.5, 0.33}, resulting in a total of 17*17*5 anchor boxes. Calculate the IOU between the anchor boxes and the ground true box to generate positive and negative samples and generate classification labels. The judgment rules are as follows:
[0040]
[0041] where y i represents the label value of the i-th sample. When the value is 1, it is a positive sample; when the value is 0, it is a negative sample; when the value is -1, this anchor box is ignored.
[0042] D5. Generate regression labels for positive samples to predict the position of the target bounding box. The label calculation method is as follows:
[0043]
[0044] where δ[i] represents the coordinate offset based on the four directions of the anchor box, T x , T y represents the center point coordinates of the target rectangle box, T w , T h represents the width and height of the target rectangle box, A x , A y represents the center point coordinates of the predefined anchor box, A w , A h represents the width and height of the predefined anchor box.
[0045] Furthermore, the optimization of the loss function in step 4 specifically includes the following steps:
[0046] E1. In the classification head, the 2k-dimensional score map obtained and the classification labels calculate the loss through the binary cross-entropy loss function. The formula is as follows:
[0047] L cls (p i , y i ) = -y i logp i - (1 - y i ) log(1 - p i ) #(10)
[0048] y i represents the label value of the i-th sample, and p i represents the classification probability value of whether the i-th sample predicted by the network belongs to the target or the background;
[0049] E2. In the regression head, the 4k-dimensional coordinate offset obtained and the regression labels calculate the loss through the SIOU loss function. The formula is as follows:
[0050]
[0051] where △ represents the distance cost between the predicted box and the ground truth box, and Ω represents the shape cost between the predicted box and the ground truth box. The specific expressions of both are as follows:
[0052]
[0053]
[0054] E3. After obtaining the classification loss and the regression loss through steps E1 and E2, the total loss of the model is shown in the following formula:
[0055] L = W box L box + W cls L cls #(14)
[0056] Finally, the model is continuously optimized through the stochastic gradient descent optimization algorithm. After 50 rounds of iterative training, the final tracking model is obtained.
[0057] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the single-object tracking method of the anti-background interference Siamese network with hierarchical feature fusion as described in any one of the above.
[0058] A non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the anti-background-interference Siamese network single-object tracking method with hierarchical feature fusion as described in any one of the above.
[0059] A computer program product including a computer program, wherein when the computer program is executed by a processor, it implements the anti-background-interference Siamese network single-object tracking method with hierarchical feature fusion as described in any one of the above.
[0060] The advantages and beneficial effects of the present invention are as follows:
[0061] 1. In a complex background, an object may be affected by occlusion, illumination changes, etc., making it difficult for single-level features to accurately describe the object. Multi-level feature fusion can combine low-level detailed information and high-level semantic information. The use of low-level features can obtain local detailed information, while the use of high-level features can obtain global semantic information. By comprehensively using this information, a more comprehensive and accurate object description can be achieved, adapting to different scales and providing richer context information, and the tracking problem in a complex background can be solved. Therefore, the present invention designs a dual-feature fusion module to fully fuse the rich spatial information captured by the lower layer of the existing network and the encoded object-level knowledge captured by the upper layer, enhancing the model's ability to use effective information to process similar background interference.
[0062] 2. Since there may be a semantic gap between different-level features, the fused feature map will inevitably introduce noise, thereby reducing the tracking accuracy. To solve this problem, context attention can be used to calculate the similarity between features and generate appropriate weights to adjust the importance of features. In this way, the model can pay more attention to the features related to the target and reduce the influence of features related to the background or noise on tracking. Therefore, the present invention designs cyclic cross attention, and through two iterative operations, each pixel can capture the long-range dependence relationship with other pixels. Such global perception helps the tracker better understand the meaning of pixels and improve the ability to solve visual understanding problems.
[0063] 3. Considering that when there is a non-overlap situation between the real frame and the predicted frame in the existing tracker, the tracking model may experience gradient disappearance, resulting in the problem of unable to accurately locate the target. In the present invention, the SIOU loss function is introduced, enabling our tracker to have the concept of intersection over union in the loss function. By redefining the penalty metric of the model, this loss function helps to obtain accurate position information during the regression process, solving the above problem. By introducing a new regression loss function, the convergence of the model is greatly optimized during the training process, and it can better measure the spatial correlation between the predicted bounding box and the actual bounding box. Thus, the tracking performance of the model in the face of a complex background is improved.
[0064] 4. By fusing semantic features and structural features, the present invention provides a more comprehensive image representation. Considering the problem that different-level feature fusion introduces noise due to the semantic gap, an attention mechanism is adopted to make the model pay more attention to the tracking target. Further, the smooth_L1 loss function used in the traditional object tracking network is replaced with a novel SIOU function to better measure the spatial correlation between the predicted bounding box and the actual bounding box, thereby improving the overall performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 FIG. is the framework diagram of the preferred embodiment BASNet anti-background interference Siamese network provided by the present invention;
[0066] Figure 2 FIG. is an example of the design process of the attention focus module;
[0067] Figure 3 FIG. is an example of the design process of the dual feature fusion module;
[0068] Figure 4 FIG. is an example of complex background tracking in the OTB100 dataset;
[0069] Figure 5 FIG. is a diagram showing the tracking effect of the present invention and the current state-of-the-art tracker. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.
[0071] The technical solution for the present invention to solve the above technical problems is as follows:
[0072] A Figure 1 FIG. is a schematic diagram of the overall network model structure of the present invention. The present invention designs a single-object tracking method for an anti-background interference Siamese network with hierarchical feature fusion - BASNet (Background-aware Hierarchical Feature Fusion Siamese Network for Visual Tracking, BASNet). In the traditional Siamese tracking network structure, a fusion module and an attention module are innovatively added to enhance the network's feature extraction ability, and the loss function is modified according to the characteristics of the existing network to further optimize the model performance. The design process of its network framework is as follows:
[0073] A single-object tracking method for an anti-background interference Siamese network with hierarchical feature fusion includes the following steps:
[0074] Step 1: Generate training data: Scale multiple datasets to obtain a template image with a pixel size of 127×127 and a search image with a pixel size of 255×255;
[0075] Step 2: Generate image features: The generated template image and search image pass through the feature extraction network AlexNet in the present invention. The second-layer low-level features and the last-layer high-level features output by the network are fused through a dual-feature fusion module to obtain multi-scale features that are beneficial to anti-background interference. And the noise introduced by the fusion of features at different layers is eliminated through the focus attention module, enabling the feature map to better focus on the target and improving the accuracy of target recognition and tracking in complex backgrounds;
[0076] Step 3: Generate training labels: After performing cross-correlation operations on the search features and the template features, a 17*17 feature response map is obtained. The classification branch determines the label value according to the IOU threshold, and the regression branch generates the label value according to the anchor box coordinate offset. Through the generated label model, the position, scale, and category information of the target can be learned, and predictions and inferences can be made based on this information;
[0077] Step 4: Loss function optimization: Input the preprocessed template features and search features into the regression branch and the search branch. Using the template features as convolution kernels, cross-correlation operations are performed on the search area to obtain their respective feature maps. The foreground and background scores on the similarity map of the classification branch and the classification label of the target are calculated using the binary cross-entropy function to obtain the loss. The regression branch uses the regression offset and the regression label of the target to calculate the loss through the new SIOU function. Finally, the loss is minimized through the stochastic gradient descent algorithm to achieve model optimization. After 50 rounds of iteration, the best tracking model is selected from them.
[0078] Preferably, in Step 1 of generating training data: Scale multiple datasets to obtain a template image with a pixel size of 127×127 and a search image with a pixel size of 255×255, which specifically includes the following steps: A1. Obtain image pairs containing object annotation information from the training dataset. The image pairs obtained from one sequence are positive sample pairs, which are used to help the model learn the appearance representation and position changes of the target; The image pairs obtained from different sequences are negative sample pairs, which are beneficial for the model to learn how to distinguish the target from the background or other objects;
[0079] A2. Since the input image sizes required by the feature extraction network are 127*127*3 and 255*255*3 respectively, it is necessary to perform central scaling or padding operations on the images. The formula is as follows:
[0080] s(w + 2p)×s(h + 2p) = A#(1)
[0081]
[0082] Where A represents the size of the scaled image, which is 127 when generating the template image and 255 when generating the search template image, w represents the image width, h represents the image height, s represents the scaling factor, and p represents the padding area. Through this operation, the information of the original image can be retained as much as possible.
[0083] Preferably, the generating of the image features in step 2 can focus on the target area during the training process, reducing the interference of background noise. By learning the differences between the target and the background, the model can better distinguish the target from the background, thereby improving the resistance to background interference. The specific process is as follows:
[0084] B1. Input the preprocessed template image and search image into the feature extraction network AlexNet. After five layers of convolution, the sizes of the template feature map and the search feature map are 6*6*256 and 22*22*256 respectively, where 6 and 22 represent the size of the feature map, and 256 represents the number of channels.
[0085] B2. In order to make the output of the last layer of the feature extraction network have the same size as the output of the second layer, deconvolution operation is performed on the last layer feature map. The last layer feature F5 output by the feature network will pass through the deconvolution layer, and the calculation formula of the deconvolution layer is as follows:
[0086] F′5=s(F5 - 1)+2p - k + 2#(5)
[0087] s = 1 represents the stride, p = 3 represents the padding margin, k = 1 represents the convolution kernel size, F′5 represents the size of the output feature map after passing through the convolution layer, F5 represents the size of the feature map input to the convolution layer. Thus, the 6*6*256 feature map becomes 12*12*256 in size.
[0088] B3. Then, add the F′5 feature map obtained in step C2 and the feature map F2 output by the second layer of the backbone network pixel by pixel to obtain F d , and the specific operation is as follows:
[0089]
[0090] Low-level features provide rich detail information and background discrimination ability, while high-level features provide semantic information of the target and shape change features.
[0091] Preferably, the attention focus module for suppressing the feature response of irrelevant regions in step 2 to reduce the influence of background interference includes the following steps:
[0092] C1. Due to the possible semantic gap between features of different layers, the fused feature map will inevitably introduce noise, thus reducing the tracking accuracy. To solve this problem, context attention can be used to calculate the similarity between features and generate appropriate weights to adjust the importance of features. The feature F obtained in step B3 is d input into the attention focus module to obtain F d′ feature. The specific formula is as follows:
[0093]
[0094] where represents the per-pixel multiplication operation, σ represents the sigmoid function, and F RCCA () represents the cyclic attention operation. Through this attention focus module, each pixel can better represent the target information, enabling the tracker to focus more on the significant features of the target and further improving the ability to solve the visual understanding problem.
[0095] C2. After obtaining the feature in step C1, it passes through a convolutional layer to reduce the dimension of the input data, reduce the amount of computation, and improve the processing speed. The convolution operation formula is as follows:
[0096]
[0097] where d w and d h represent the width and height of the input data matrix, k w and k h represent the width and height of the convolution kernel, s represents the moving stride, and p represents the input data padding value. The convolution operation, through the sliding window method of the convolution kernel, can model the spatial relationship in the image. This helps the model capture the context information in the area around the target, thus better distinguishing the target from the background.
[0098] Preferably, generating the training label in step 3 is beneficial to providing annotation information about the target position, scale, and category and is used for the training of the model, thereby improving the accuracy and robustness of target tracking. The specific steps are as follows:
[0099] D1. The feature dimensions obtained in step C2 are respectively input into the classification head and the regression head, and 4*4*256 and 20*20*256 are obtained through a convolution respectively. The purpose of this operation is to transform the template feature into a low-dimensional representation suitable for the classifier, which can improve the computational efficiency while retaining the key information and maintaining an appropriate local response ability.
[0100] D2. The template images of the input classification branch and regression branch are upsampled by a convolutional kernel, from 4*4*256 to 4*4*(2k*256) and 4*4*(4k*256). The purpose of this upsampling operation is to make the template features between different targets have a large separation in the feature space, so as to better distinguish the feature patterns of different targets. The search features only need to perform a correlation calculation with the template features, without directly representing the appearance information of the target, so there is no need for upsampling.
[0101] D3. Use the template image features as a convolutional kernel to perform a cross-correlation operation with the search image features. The similarity map obtained by the classification branch is 17*17*2k, and the similarity map obtained by the regression branch is 17*17*4k, where k refers to the number of anchor boxes.
[0102] D4. Generate five anchor boxes with different aspect ratios for each pixel of the 17*17 feature map, and the scale distribution is {3, 2, 1, 0.5, 0.33}, a total of 17*17*5 anchor boxes. Calculate the IOU between the anchor boxes and the ground true box to generate positive and negative samples. Generate classification labels, and the judgment rules are as follows:
[0103]
[0104] Among them, y i represents the label value of the i-th sample. When the value is 1, it is a positive sample; when the value is 0, it is a negative sample; when the value is -1, this anchor box is ignored.
[0105] D5. Generate regression labels for positive samples to predict the position of the target bounding box. The label calculation method is as follows:
[0106]
[0107] Among them, δ[i] represents the coordinate offset based on the four directions of the anchor box, T x , T y represents the center point coordinates of the target rectangle box, T w , T h represents the width and height of the target rectangle box, A x , A y represents the center point coordinates of the predefined anchor box, A w , A h represents the width and height of the predefined anchor box.
[0108] Preferably, the optimization of the loss function in step 4 can accelerate the model convergence speed and improve the model performance. The specific steps are as follows:
[0109] E1. In the classification head, the loss is calculated between the obtained 2k-dimensional score map and the classification labels using the binary cross-entropy loss function, and the formula is as follows:
[0110] L cls (p i , y i ) = -y i log p i - (1 - y i ) log(1 - p i ) #(10)
[0111] y i represents the label value of the i-th sample, and p i represents the classification probability value of whether the i-th sample predicted by the network belongs to the target or the background.
[0112] E2. In the regression head, the loss is calculated between the obtained 4k-dimensional coordinate offset and the regression labels using the SIOU loss function, and the formula is as follows:
[0113]
[0114] where Δ represents the distance cost between the predicted box and the ground truth box. The larger Δ is, the greater the gap between the predicted value and the true value. Ω represents the shape cost between the predicted box and the ground truth box. The larger Ω is, the greater the shape difference between the predicted box and the ground truth box. The specific expressions of the two are as follows:
[0115]
[0116]
[0117] E3. After obtaining the classification loss and the regression loss through steps E1 and E2, the total loss of the model is shown in the following formula:
[0118] L = W box L box + W cls L cls #(14)
[0119] Finally, the model is continuously optimized through the stochastic gradient descent optimization algorithm. After 50 rounds of iterative training, the final tracking model is obtained.
[0120] The present invention proposes a BASNet network to solve the problem of tracking failure in complex environments due to relying on single-layer features. Specifically, first, a dual-feature fusion module is proposed, which applies the outputs of the second layer and the fifth layer of the network together. This is because the combination of low-level features and high-level features can bring out the advantages of different levels while obtaining more target features, which is more conducive to accurately locating the target. Secondly, the present invention designs an attention focus module. By considering the environmental features around the target and the positional relationship and shape of other objects, it helps the tracking algorithm to more accurately determine the position and boundary of the target, reducing the possibility of misjudging objects in the background as the target. Finally, a new loss function, SIoU, is introduced. This loss function can encourage the tracking algorithm to better adapt to complex backgrounds and reduce the overlapping part between the target and the background. The present invention improves the tracking performance of the target tracker in complex background situations.
[0121] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0122] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0123] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0124] The above embodiments should be understood as being only for the purpose of illustrating the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A hierarchical feature fusion and background interference resistant twin network single target tracking method, characterized in that: The following steps are involved: Step 1: Generate training data: Scale multiple data sets to obtain a template image with a pixel size of 127×127 and a search image with a pixel size of 255×255; Step 2: Generate image features: The generated template image and search image are passed through the feature extraction network AlexNet. The second layer of low-level features and the last layer of high-level features output by the network are fused through the dual feature fusion module to obtain multi-scale features. The noise introduced by the fusion of different layers of features is eliminated through the attention focus module. Step 3: Generate training labels: After cross-correlating the search features with the template features, a 17*17 feature response map is obtained. The classification branch determines the label value according to the IOU threshold, and the regression branch generates the label value according to the anchor box coordinate offset. The generated label model is used to learn the location, scale, and category information of the target, and prediction and reasoning are performed based on this information. Step 4, loss function optimization: input the preprocessed template features and search features into the regression branch and search branch; use the template features as the convolution kernel, perform cross-correlation operations on the search area to obtain their respective feature maps, use the binary cross entropy function to calculate the loss of the foreground and background scores on the similarity map of the classification branch and the classification label of the target, and use the regression offset and the regression label of the target in the regression branch. The regression branch calculates the loss through the new SIOU function; finally, the stochastic gradient descent algorithm is used to minimize the loss to achieve model optimization. After several rounds of iterations, the best tracking model is selected; In step 3, training labels are generated. The specific steps are as follows: D1, input the feature dimensions obtained in step C2 into the classification head and regression head respectively, and obtain 4*4*256 and 20*20*256 respectively after a convolution. This operation converts the template features into a low-dimensional representation suitable for the classifier; D2. The template image of the input classification branch and regression branch is upgraded through a convolution kernel from 4*4*256 to 4*4*(2k*256) and 4*4*(4k*256). This dimensionality increase operation is used to distinguish the feature patterns of different targets; the search feature is not upgraded; D3, the template image features are used as convolution kernels to perform cross-correlation operations with the search image features. The similarity map obtained by the classification branch is 17*17*2k, and the similarity map obtained by the regression branch is 17*17*4k, where k refers to the number of anchor boxes; D4. Generate five anchors with different aspect ratios for each pixel of the 17*17 feature map. The scales are {3, 2, 1, 0.5, 0.33}, for a total of 17*17*5 anchors. Calculate the IOU between the anchor and the ground true box to generate positive and negative samples, and generate classification labels. The judgment rules are as follows: where y i Indicates the label value of the i-th sample. When the value is 1, it is a positive sample, when the value is 0, it is a negative sample, and when the value is -1, this anchor box is ignored; D5. Generate regression labels for positive samples to predict the bounding box position of the target. The label calculation method is as follows: Where δ[i] represents the coordinate offset based on the four directions of the anchor box, T x , T y Indicates the coordinates of the center point of the target rectangle, T w , T h Indicates the width and height of the target rectangle, A x , A y Represents the center point coordinates of the predefined anchor box, A w , A h Represents the width and height of the predefined anchor box; The loss function optimization in step 4 specifically includes the following steps: E1. In the classification head, the obtained 2k-dimensional score map and the classification label are calculated through the binary cross entropy loss function. The formula is as follows: L cls (p i ,y i )=-y i logp i -(1-y i )log(1-p i )#(10) y i represents the label value of the i-th sample, p i Indicates the classification probability value of whether the i-th sample predicted by the network belongs to the target or the background; E2. In the regression head, the obtained 4k-dimensional coordinate offset and regression label are used to calculate the loss through the SIOU loss function. The formula is as follows: Among them, △ represents the distance cost between the predicted box and the real box, and Ω represents the shape cost between the predicted box and the real box. The specific expressions of the two are as follows: E3. After obtaining the classification loss and regression loss through steps E1 and E2, the total loss of the model is shown in the following formula: L=W box L box +W cls L cls #(14) Finally, the model is continuously optimized through the stochastic gradient descent optimization algorithm, and the final tracking model is obtained after 50 rounds of iterative training.
2. According to claim 1, the hierarchical feature fusion and background interference resistant twin network single target tracking method is characterized in that: The step 1 of generating training data specifically includes the following steps: A1. Obtain image pairs containing object annotation information from the training data set. Image pairs obtained from one sequence are positive sample pairs, which are used to help the model learn the appearance representation and position changes of the target; image pairs obtained from different sequences are negative sample pairs, which are used for the model to learn how to distinguish between the target and the background or other objects. A2. Perform center scaling or filling operations on the image. The formula is as follows: s(w+2p)×s(h+2p)=A#(1) A represents the scaled image size, which is 127 when generating a template image and 255 when generating a search template image. w represents the image width, h represents the image height, s represents the scaling factor, and p represents the padding area.
3. According to claim 1, the hierarchical feature fusion and background interference resistant twin network single target tracking method is characterized in that: The step 2 generates image features, and the specific process is as follows: B1. Input the preprocessed template image and search image into the feature extraction network AlexNet. After five layers of convolution, the template feature map size and search feature map size are 6*6*256 and 22*22*256 respectively, where 6 and 22 represent the feature map size, and 256 represents the number of channels. B2. In order to make the last layer output of the feature extraction network have the same size as the second layer output, the deconvolution operation is performed on the last layer feature map. The last layer feature F5 output by the feature network will pass through the deconvolution layer. The calculation formula of the deconvolution layer is as follows: F′5=s(F5-1)+2p-k+2#(5) s=1 represents the step size, p=3 represents the padding margin, k=1 represents the convolution kernel size, F′5 represents the output feature map size after the convolution layer, and F5 represents the feature map size of the input convolution layer. Thus, the 6*6*256 feature map becomes 12*12*256 in size. B3, then add the F′5 feature map obtained in step C2 to the feature map F2 output by the second layer of the backbone network pixel by pixel to obtain F d , the specific operations are as follows: Low-level features provide rich detail information and background distinction capabilities, while high-level features provide semantic information and shape change characteristics of the target.
4. According to claim 3, the hierarchical feature fusion and background interference resistant twin network single target tracking method is characterized in that: The attention focus module in step 2 includes the following steps: C1, context attention is used to calculate the similarity between features and generate weights to adjust the importance of features; the feature F obtained in step B3 is d Input into the attention focus module to get F d′ Features, the specific formula is as follows: in represents the pixel-by-pixel multiplication operation, σ represents the sigmoid function, and F RCCA () represents the cyclic attention operation. Through this attention focus module, each pixel can better represent the target information, so that the tracker can focus more on the salient features of the target; C2. After the features are obtained in step C1, a convolution layer is passed to reduce the dimension of the input data. The convolution operation formula is as follows: where d w d h Represents the width and height of the input data matrix, k w , k h Represents the width and height of the convolution kernel, s represents the moving stride, and p represents the input data padding value; the convolution operation can model the spatial relationship in the image through the sliding window method of the convolution kernel; The model is used to capture the contextual information of the area around the target, thereby distinguishing the target from the background.
5. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for single target tracking by a twin network with hierarchical feature fusion and background interference resistance as described in any one of claims 1 to 4 is implemented.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the background interference resistant twin network single target tracking method with hierarchical feature fusion as described in any one of claims 1 to 4.
7. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the background interference resistant twin network single target tracking method with hierarchical feature fusion as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Optical flow calculation method combining image pyramid guidance and cyclic cross attention
CN114821105A