Foreground Information Guided Siamese Convolutional Neural Network Target Tracking Method and Device

Through improved residual network and sampling strategy, the positive sample pairs are enriched, combined with the fill loss function, the recognition ability and robustness of the twin convolutional neural network target tracking method in complex backgrounds is enhanced, and the anti-interference problem of existing methods under factors such as occlusion and deformation is solved.

CN114693728BActive Publication Date: 2025-07-18CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011610435.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-30
Publication Date
2025-07-18
Estimated Expiration
2040-12-30

AI Technical Summary

Technical Problem

The existing target tracking method based on twin convolutional neural networks is insufficient in complex scenarios, especially in the absence of robustness and accuracy under factors such as occlusion and deformation.

Method used

The improved residual network is used to obtain the deep features of higher semantic information, increase the challenge factors in the positive sample pair through the sampling strategy, and introduce fill loss calculation methods into the loss function to enhance the significance of the foreground information and reduce background interference.

Benefits of technology

It improves the recognition ability and robustness of the target tracking method in complex backgrounds, and reduces the impact of background interference on tracking performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693728B_ABST
    Figure CN114693728B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of planning, and particularly relates to a foreground information-guided twin convolutional neural network object tracking method and device. The method and device use an improved residual network to replace the shallow network to obtain depth features with higher semantic information content; adopt a sampling strategy to enhance the challenging factors in positive sample pairs; input the template image, the search region image, and the padding image into the template branch, the search region branch, and the guidance branch respectively to extract features, and perform depth cross-correlation calculation among the feature tensors output by each branch. In the calculation, the foreground information is used to improve the loss function by means of the padding loss calculation method; the padding loss is introduced based on the improved loss function, and the network model is guided to complete training. The method and device at least solve the technical problem of weak anti-interference factors in existing object tracking methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and in particular, to a foreground information-guided Siamese convolutional neural network target tracking method and device. Background Art

[0002] As one of the most important research directions in the field of computer vision, target tracking has always received extensive attention from scholars. Given the bounding box information of any object in an image, the purpose of a target tracking method is to accurately locate it in subsequent images. Although target tracking methods have been greatly improved in recent years and are widely used in many advanced fields, due to the influence of factors such as occlusion, illumination change, and similar background interference in complex scenes, the target tracking task still poses challenges.

[0003] In recent years, target tracking methods based on Siamese convolutional neural networks have received more attention due to their balanced accuracy, robustness, and real-time performance. Such methods achieve precise matching between deep features by learning an accurate and robust similarity metric function, thereby accurately and effectively locating the target. Among them, the SiamFC tracking method proposed by Bertinetto et al. is particularly remarkable. It first uses a fully convolutional network to extract the deep features of the template and the search region, and applies cross-correlation calculation to the output deep features to obtain the response map after matching at one time. This method not only performs excellently in tracking performance but also perfectly meets the real-time requirement in terms of running speed. Subsequently, many scholars have improved and innovated based on this method. Valmadre et al. introduced a correlation filter into the method and introduced a differentiable layer in the similarity metric calculation of SiamFC. Guo et al. introduced a template update mechanism into the method and updated the template features by learning the changes in the background and the appearance of the target, thereby improving the robustness and accuracy of the method. Li et al. introduced the RPN (Region Proposal Network) model commonly used in target detection methods into the method, which can not only effectively handle the multi-scale changes of the target but also greatly improve the accuracy and robustness of the method. Zhu et al. introduced an interference perception module based on the research of Li et al. and effectively expanded the training sample set, successfully improving the robustness of the algorithm in complex scenes. Based on previous research, Li et al. successfully introduced a deep network architecture into the method, which not only surpassed most of the previous methods but also broke through the limitations of the shallow network architecture, opening up a new research trend.

[0004] Although the above tracking methods have made great progress in tracking performance, there are still some unimproved defects in the object tracking method based on the Siamese convolutional neural network. First, in the offline training stage, the challenging factors (occlusion, deformation, etc.) contained in the positive sample pairs used for network training are not rich enough. Most object tracking methods based on the Siamese convolutional neural network will use datasets containing rich categories and training data for training, such as the VID dataset of ILSVRC-2015. However, these datasets with limited categories are not sufficient to provide sufficient positive sample pairs, making it difficult to train a network model with high quality and strong robustness. Even though the DaSiamRPN method expands the number of positive sample pairs by introducing other large-scale training datasets, due to the long-tail distribution of some interference factors (occlusion, deformation, etc.), the expanded positive sample pairs still cannot contain rich enough interference factors to improve the robustness of the method. Second, when the background is relatively complex, the object tracking method based on the Siamese convolutional neural network cannot maintain stable tracking performance. Most object tracking methods based on the Siamese convolutional neural network can locate the tracking object from some simple non-semantic backgrounds. However, when encountering some complex scenes during the tracking process, the semantic background, which is considered the main interference source, is the key factor affecting the tracking performance. When the background is relatively complex, the tracking box will gradually drift to the interference sources in the background, resulting in unstable tracking or even tracking failure. Summary of the Invention

[0005] Embodiments of the present invention provide a foreground information-guided Siamese convolutional neural network object tracking method and device to at least solve the technical problem of weak anti-interference factors in existing object tracking methods.

[0006] According to an embodiment of the present invention, a foreground information-guided Siamese convolutional neural network object tracking method is provided, including the following steps:

[0007] Replace the shallow network with an improved residual network to obtain depth features with higher semantic information content;

[0008] Adopt a sampling strategy to increase the challenging factors in the positive sample pairs;

[0009] Input the template image, the search region image, and the padding image into the template branch, the search region branch, and the guiding branch respectively to extract features, and perform depth cross-correlation calculations among the feature tensors output by each branch. In the calculation, use the padding loss calculation method for the foreground information to improve the loss function;

[0010] Based on the improved loss function, introduce the padding loss and guide the network model to complete training.

[0011] Further, the method further includes the step:

[0012] In the tracking stage, the template image and the search region image are respectively input into the template branch and the search region branch to extract features;

[0013] Cross-correlation calculation is performed using the feature tensors output by the two branches to obtain the final response map;

[0014] Find the maximum peak in the obtained response map and map it to the search region image to determine the exact position of the target.

[0015] Furthermore, replacing the shallow network with the improved residual network includes:

[0016] In the improved network, the stride of the last two convolutional blocks in the residual network is changed to 1, and a convolutional layer with a kernel size of 1×1 is added at the end of the network to reduce the output dimension.

[0017] Furthermore, adopting a sampling strategy to enhance the challenging factors in positive sample pairs includes:

[0018] The sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference;

[0019] The sampling strategy uses a random affine transformation combining rotation and shearing mapping to simulate the deformation of the target.

[0020] Furthermore, the sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference includes:

[0021] In the offline training stage of the network model, the sampling strategy regards the template image in the positive sample pair as the object to be augmented; the actual size W×H of the target is obtained according to the target information in the template image; after obtaining the target size, an occlusion mask with a pre-designed direction is generated according to this information, and the size of the mask is fixed as W / 2×H / 2; at the same time, each occlusion mask removes the image pixels it covers, so as to simulate the occlusion interference in the tracking process and improve the robustness when dealing with occlusion.

[0022] Furthermore, the sampling strategy uses a random affine transformation combining rotation and shearing mapping to simulate the deformation of the target includes:

[0023] When a template image is given, the sampling strategy rotates the template image within an angle range of θ = ±30°; at the same time, the template image is subjected to shearing mapping in the X and Y directions, and the mapping angle ranges in the two directions are φ = ±25° and ψ = ±25° respectively; the above two transformations are randomly combined to form the random affine transformation adopted in the sampling strategy.

[0024] Furthermore, in the calculation, using the filling loss calculation method for foreground information to improve the loss function includes:

[0025] Make full use of the foreground information and use the filling loss calculation method to improve the loss function of the method. The input of the guiding branch is a filled image with a size of 255×255×3; the filled image is obtained by filling the background with the mean value of the entire image, and the filling method is as follows:

[0026]

[0027] In the formula, a and b represent the a-th row and the b-th column in the image, fg represents the foreground information, and bg represents the background information; and the network model used for feature extraction in the guiding branch is the same as the models in the target template branch and the search area branch, and the improved ResNet-18 network is used as the feature extraction network;

[0028] When the template image, the search area image, and the filled image are given, each image is input into the corresponding branch to extract convolutional features; then, deep cross-correlation calculation is performed among the feature tensors output by each branch; among them, the deep cross-correlation result between the target template branch and the search area branch is set as S, and the deep cross-correlation result between the target template branch and the guiding branch is S'.

[0029] Furthermore, the filling loss introduced on the basis of the improved loss function includes:

[0030] After performing the deep cross-correlation calculation among the branches, the filling loss is introduced on the basis of the loss function of the SiamFC method, and the calculation method of the filling loss is as follows:

[0031]

[0032] In the formula, C, H, and W respectively represent the number of channels, the height, and the width of the deep cross-correlation result; the loss function after fusing the filling loss is defined as:

[0033] loss final =(1 - λ)loss training +λloss padding (1.3)

[0034] In the formula, λ is the weight for fusing the calculations of each loss.

[0035] According to another embodiment of the present invention, a foreground information-guided siamese convolutional neural network target tracking device is provided, including:

[0036] A network replacement unit for replacing the shallow network with an improved residual network to obtain deeper features with higher semantic information content;

[0037] A factor improvement unit for using a sampling strategy to improve the challenging factors in the positive sample pair;

[0038] A cross - correlation calculation unit, configured to input the template image, the search region image, and the padding image into the template branch, the search region branch, and the guiding branch respectively to extract features, and perform depth cross - correlation calculation among the feature tensors output by each branch, and use a padding loss calculation method for foreground information in the calculation to improve the loss function;

[0039] A loss padding unit, configured to reference the padding loss based on the improved loss function and guide the network model to complete training.

[0040] Furthermore, the device further includes:

[0041] A feature extraction unit, configured to input the template image and the search region image into the template branch and the search region branch respectively to extract features during the tracking stage;

[0042] A response map acquisition unit, configured to perform cross - correlation calculation using the feature tensors output by the two branches to obtain the final response map;

[0043] A position localization unit, configured to find the maximum peak in the obtained response map and map it to the search region image to determine the exact position of the target.

[0044] In the foreground information - guided Siamese convolutional neural network target tracking method and device according to the embodiments of the present invention, a relatively simple sampling strategy is adopted to expand the challenging factors in the positive sample pairs during offline training. And in order to further improve the recognition ability of the method in the semantic background, the present invention adopts the method of filling and occluding the background information to enhance the saliency of the foreground information. At the same time, the processed data is input into the guiding branch based on the convolutional neural network, and a padding loss calculation method is used to improve the loss function of the method, effectively improving the recognition ability of the method in the semantic background. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0046] Figure 1 is a schematic diagram of the technical process of the present invention;

[0047] Figure 2 is a schematic diagram of the tracking process provided by the embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of the occlusion mask provided by the embodiment of the present invention;

[0049] Figure 4Schematic diagram of random affine transformation provided by an embodiment of the present invention. Detailed implementation manners

[0050] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] Embodiment 1

[0053] According to an embodiment of the present invention, a foreground information-guided twin convolutional neural network object tracking method is provided. Refer to Figure 1 , and it includes the following steps:

[0054] Use the improved residual network to replace the shallow network to obtain depth features with higher semantic information content;

[0055] Adopt a sampling strategy to improve the challenging factors in the positive sample pair;

[0056] Input the template image, the search region image, and the padding image into the template branch, the search region branch, and the guidance branch respectively to extract features, and perform depth cross-correlation calculation among the feature tensors output by each branch. In the calculation, use the padding loss calculation method for the foreground information to improve the loss function;

[0057] Based on the improved loss function, introduce the padding loss and guide the network model to complete training.

[0058] In the foreground information-guided twin convolutional neural network target tracking method in the embodiments of the present invention, a relatively simple sampling strategy is adopted to augment the challenging factors in the positive sample pairs during offline training. And in order to further improve the recognition ability of this method in the semantic background, the present invention adopts the method of filling and occluding the background information to enhance the saliency of the foreground information. At the same time, the processed data is input into the guiding branch based on the convolutional neural network, and a filling loss calculation method is used to improve the loss function of the method, effectively improving the recognition ability of the method in the semantic background.

[0059] Among them, the method further includes the steps of:

[0060] In the tracking stage, the template image and the search region image are respectively input into the template branch and the search region branch to extract features;

[0061] Cross-correlation calculation is performed using the feature tensors output by the two branches to obtain the final response map;

[0062] Find the maximum peak in the obtained response map and map it to the search region image to determine the exact position of the target.

[0063] Among them, replacing the shallow network with the improved residual network includes:

[0064] In the improved network, the stride of the last two convolutional blocks in the residual network is changed to 1, and a convolutional layer with a kernel size of 1×1 is added at the end of the network to reduce the output dimension.

[0065] Among them, adopting a sampling strategy to increase the challenging factors in the positive sample pairs includes:

[0066] The sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference;

[0067] The sampling strategy adopts a combination of rotation and shear mapping random affine transformation to simulate the deformation of the target.

[0068] Among them, the sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference includes:

[0069] In the offline training stage of the network model, the sampling strategy regards the template image in the positive sample pair as the object to be augmented; obtains the actual size W×H of the target according to the target information in the template image; after obtaining the target size, generates an occlusion mask with a pre-designed direction according to this information, and fixes the size of the mask to W / 2×H / 2; at the same time, each occlusion mask removes the image pixels it covers, so as to simulate the occlusion interference in the tracking process and improve the robustness when dealing with occlusion.

[0070] Among them, the sampling strategy uses a random affine transformation that combines rotation and shear mapping to simulate the deformation of the target, including:

[0071] When a template image is given, the sampling strategy rotates the template image within an angle range of θ = ±30°; at the same time, the template image will be subjected to shear mapping in the X and Y directions, and the mapping angle ranges in the two directions are φ = ±25° and ψ = ±25° respectively; the above two transformations are randomly combined to form the random affine transformation adopted in the sampling strategy.

[0072] Among them, in the calculation, the foreground information is used to improve the loss function by using the filling loss calculation method, including:

[0073] The foreground information is fully utilized and the filling loss calculation method is used to improve the loss function of the method. The input of the guiding branch is a filled image with a size of 255×255×3; this filled image is obtained by filling the background with the mean value of the entire image, and the filling method is:

[0074]

[0075] In the formula, a and b represent the a-th row and the b-th column in the image, fg represents the foreground information, and bg represents the background information; and the network model used for feature extraction in the guiding branch is the same as the models in the target template branch and the search area branch, and the improved ResNet-18 network is used as the feature extraction network.

[0076] When the template image, the search area image, and the filled image are given, each image is input into the corresponding branch to extract convolutional features; then, depth cross-correlation calculations are performed between the feature tensors output by each branch; among them, the depth cross-correlation result between the target template branch and the search area branch is set as S, and the depth cross-correlation result between the target template branch and the guiding branch is S'.

[0077] Among them, the filling loss is introduced based on the improved loss function, including:

[0078] After performing the depth cross-correlation calculations between the branches, the filling loss is introduced based on the loss function of the SiamFC method, and the calculation method of the filling loss is:

[0079]

[0080] In the formula, C, H, and W respectively represent the number of channels, the height, and the width of the depth cross-correlation result; the loss function after fusing the filling loss is defined as:

[0081] loss final =(1 - λ)loss training +λloss padding(1.3)

[0082] where λ is the weight for fusing the loss calculations.

[0083] The following is a detailed description of the foreground information-guided Siamese convolutional neural network target tracking method of the present invention with specific embodiments:

[0084] Aiming at the problem that the existing target tracking method based on the Siamese convolutional neural network has a relatively shallow network layer, and the tracking performance is unstable or even loses the tracking target when the target is severely deformed, there is severe occlusion in the background, and there is interference from similar backgrounds, the present invention proposes a foreground information-guided Siamese convolutional neural network target tracking method. This method is based on the SiamFC method and adopts a relatively simple sampling strategy to expand the challenging factors in the positive sample pairs during offline training. And in order to further improve the recognition ability of this method in the semantic background, the present invention uses the method of filling and occluding the background information to enhance the saliency of the foreground information. At the same time, the processed data is input into the guiding branch based on the convolutional neural network, and a filling loss calculation method is used to improve the loss function of the method, effectively improving the recognition ability of the method in the semantic background.

[0085] See Figures 1-4 , the specific solution of this foreground information-guided Siamese convolutional neural network target tracking method is as follows:

[0086] Use the improved residual network to replace the shallow network to obtain deeper features with richer semantic information. In the improved network, the strides of the last two convolutional blocks Layer3 and Layer4 in the residual network ResNet-18 are changed to 1, and a convolutional layer Conv2 with a kernel size of 1×1 is added at the end of the network to reduce the output dimension. The improved network structure is shown in Table 1 below:

[0087]

[0088]

[0089] Table 1 Network Structure

[0090] Among them, the data order of the convolutional layer and convolutional block (Conv and Layer) in the layer structure is kernel size, number of channels, stride, and padding number, and the data order of the pooling layer (Maxpool) is kernel size, stride, and padding number.

[0091] A simple sampling strategy is adopted to enrich the challenging factors in the positive sample pairs. First, the sampling strategy uses artificially designed occlusion masks to increase the number of positive sample pairs under occlusion interference. In the offline training stage of the network model, this strategy regards the template image in the positive sample pair as the object to be augmented. Specifically, according to the target information in the template image, the actual size W×H of the target can be obtained. After obtaining the target size, an occlusion mask with a pre-designed direction can be generated according to this information, and the size of the mask is fixed at W / 2×H / 2. At the same time, each occlusion mask removes the image pixels it covers to simulate the occlusion interference in the tracking process and improve the robustness of the method in dealing with occlusion. Second, the sampling strategy uses a combination of rotation and shear mapping (Shear Mapping) random affine transformation to simulate the deformation of the target. When a template image is given, this strategy rotates the template image within the range of θ = ±30°. At the same time, the template image will be subjected to shear mapping in the X and Y directions, and the mapping angle ranges in the two directions are φ = ±25° and ψ = ±25° respectively. The above two transformations are randomly combined to form the random affine transformation adopted in the augmentation strategy, which can not only enrich the dataset but also improve the robustness of the tracking method in dealing with target deformation.

[0092] Make full use of the foreground information and adopt a filling loss calculation method to improve the loss function of the method, so as to guide the tracking method to pay more attention to the changes in foreground information during training and effectively reduce the influence of background interference on the method. The input of the guiding branch is a filled image with a size of 255×255×3. This filled image is obtained by filling the background with the mean value of the whole image, and the filling method is shown in the following formula:

[0093]

[0094] In the formula, a and b represent the a-th row and the b-th column in the image, fg represents the foreground information, and bg represents the background information. Moreover, the network model used for feature extraction in the guiding branch is the same as that in the target template branch and the search region branch, and the improved ResNet-18 network is used as the feature extraction network (as shown in Table 1). When the template image, the search region image, and the filled image are given, the present invention inputs each image into the corresponding branch to extract convolutional features. Then, the present invention performs depthwise cross correlation (DW-XCorr) calculation between the feature tensors output by each branch. Among them, the depthwise cross correlation result between the target template branch and the search region branch is set as S, and the depthwise cross correlation result between the target template branch and the guiding branch is S'.

[0095] After performing the depth cross-correlation calculation between branches, the present invention introduces a padding loss based on the loss function of the SiamFC method, thereby guiding the method to pay more attention to the change of foreground information during training and reducing the influence of background interference on the algorithm. The calculation method of the padding loss is shown in the following formula:

[0096]

[0097] In the formula, C, H, and W respectively represent the number of channels, height, and width of the depth cross-correlation result. Using this padding loss, the method will be guided to pay less attention to background information during training, improve the attention of the algorithm to foreground information, and effectively improve the recognition ability of the algorithm in the semantic background. The loss function after fusing the padding loss can be defined as:

[0098] loss final =(1 - λ)loss training +λloss padding (1.3)

[0099] In the formula, λ is the weight for fusing the calculations of each loss. It should be noted that the foreground information guidance proposed by the present invention is mainly applied in the training stage of the network. When the network training is completed and tracking is performed, the guidance branch will be blocked.

[0100] As Figure 1 shown, it is a schematic diagram of the technical process of a foreground information-guided Siamese convolutional neural network target tracking method provided in an embodiment of the present invention. As Figure 2 shown, it is a schematic diagram of the tracking process of a foreground information-guided Siamese convolutional neural network target tracking method provided in an embodiment of the present invention.

[0101] In the training stage, first, the template image, the search region image, and the padding image are respectively input into the template branch, the search region branch, and the guidance branch to extract features. The images of each branch are data processed in advance using a sampling strategy, and the sizes are fixed at 127×127×3, 255×255×3, and 255×255×3 respectively. Then, a depth cross-correlation calculation is performed between the feature tensors output by each branch. Finally, the loss function is calculated using formulas (1.2) and (1.3), and the network model is guided to complete training.

[0102] In the tracking stage, first, the template image and the search region image are respectively input into the template branch and the search region branch to extract features. Then, a cross-correlation calculation is performed using the feature tensors output by the two branches to obtain the final response map. Finally, the maximum peak is found in the obtained response map and mapped to the search region image to determine the exact position of the target. To solve the problem of target scale change, the present invention stipulates that the mapped scale estimation is three values of 1.0375{-1,0,1}.

[0103] Example 2

[0104] According to another embodiment of the present invention, a foreground information guided twin convolutional neural network target tracking device is provided, see Figures 1-4 ,include:

[0105] The network replacement unit is used to replace the shallow network with the improved residual network to obtain deep features with higher semantic information;

[0106] A factor enhancement unit, used to adopt a sampling strategy to enhance the challenge factor in the positive sample pair;

[0107] A cross-correlation calculation unit is used to input the template image, the search area image and the filling image into the template branch, the search area branch and the guide branch respectively to extract features, and perform deep cross-correlation calculation between the feature tensors output by each branch. In the calculation, the filling loss calculation method is used for the foreground information to improve the loss function;

[0108] The loss filling unit is used to reference the filling loss based on the improved loss function and guide the network model to complete the training.

[0109] The foreground information guided twin convolutional neural network target tracking device in the embodiment of the present invention adopts a relatively simple sampling strategy to expand the challenging factors in the positive sample pairs during offline training. And in order to further improve the recognition ability of the method in the semantic context, the present invention adopts the method of filling and blocking the background information to enhance the significance of the foreground information. At the same time, the processed data is input into the guidance branch based on the convolutional neural network, and a filling loss calculation method is used to improve the loss function of the method, which effectively improves the recognition ability of the method in the semantic context.

[0110] The device further comprises:

[0111] A feature extraction unit, used to input the template image and the search area image into the template branch and the search area branch respectively to extract features during the tracking phase;

[0112] A response map acquisition unit, used to perform cross-correlation calculation using the feature tensors output by the two branches, so as to obtain a final response map;

[0113] The position positioning unit is used to find the maximum peak in the obtained response map and map it to the search area image to determine the exact position of the target.

[0114] See also Figures 1-4 , the specific scheme of the foreground information guided twin convolutional neural network target tracking device is as follows:

[0115] Network replacement unit: The improved residual network is used to replace the shallow network to obtain depth features with richer semantic information. In the improved network, the strides of the last two convolutional blocks Layer3 and Layer4 in the ResNet-18 are changed to 1, and a convolutional layer Conv2 with a kernel size of 1×1 is added at the end of the network to reduce the output dimension. The improved network structure is shown in Table 1 above.

[0116] Among them, the data order of the convolutional layer and convolutional block (Conv and Layer) in the layer structure is kernel size, number of channels, stride, and padding number, and the data order of the pooling layer (Maxpool) is kernel size, stride, and padding number.

[0117] Factor enhancement unit: A simple sampling strategy is adopted to enrich the challenging factors in the positive sample pairs. First, the sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference. In the offline training stage of the network model, the template image in the positive sample pair is regarded as the object to be augmented. Specifically, the actual size W×H of the object can be obtained according to the target information in the template image. After obtaining the target size, an occlusion mask with a pre-designed direction can be generated according to this information, and the size of the mask is fixed at W / 2×H / 2. At the same time, each occlusion mask removes the image pixels it covers to simulate the occlusion interference in the tracking process and improve the robustness of the method in dealing with occlusion. Second, the sampling strategy uses a combination of rotation and shear mapping to simulate the deformation of the target. When a template image is given, the strategy rotates the template image within an angle range of θ = ±30°. At the same time, the template image is subjected to shear mapping in the X and Y directions, and the mapping angle ranges in the two directions are φ = ±25° and ψ = ±25° respectively. The above two transformations are randomly combined to form the random affine transformation adopted in the augmentation strategy, which can not only enrich the dataset but also improve the robustness of the tracking method in dealing with target deformation.

[0118] Cross-correlation calculation unit: Make full use of the foreground information and use a filling loss calculation method to improve the loss function of the method, so as to guide the tracking method to pay more attention to the changes in foreground information during training and effectively reduce the influence of background interference on the method. The input of the guiding branch is a filled image with a size of 255×255×3. This filled image is obtained by filling the background with the mean value of the whole image, and the filling method is shown in the following formula:

[0119]

[0120] Wherein, a and b represent the a-th row and the b-th column in the image, fg represents foreground information, and bg represents background information. Moreover, the network model for feature extraction in the guiding branch is the same as that in the target template branch and the search region branch, and the improved ResNet-18 network is adopted as the feature extraction network (as shown in Table 1). When the template image, the search region image, and the padding image are given, the present invention will input each image into the corresponding branch to extract convolutional features. Then, the present invention will perform depthwise cross-correlation (DW-XCorr) calculation among the feature tensors output from each branch. Among them, the depthwise cross-correlation result between the target template branch and the search region branch is set as S, and the depthwise cross-correlation result between the target template branch and the guiding branch is S'.

[0121] Loss padding unit: After performing the depthwise cross-correlation calculation among each branch, the present invention will introduce a padding loss on the basis of the loss function of the SiamFC method, so as to guide the method to pay more attention to the change of foreground information during training and reduce the influence of background interference on the algorithm. The calculation method of the padding loss is shown in the following formula:

[0122]

[0123] Wherein, C, H, and W respectively represent the number of channels, the height, and the width of the depthwise cross-correlation result. By using this padding loss, the present method will be guided to pay less attention to background information during training, improve the attention of the algorithm to foreground information, and effectively improve the recognition ability of the algorithm in the semantic background. The loss function after fusing the padding loss can be defined as:

[0124] loss final =(1 - λ)loss training +λloss padding (1.3)

[0125] Wherein, λ is the weight for fusing the calculations of each loss. It should be noted that the foreground information guidance proposed by the present invention is mainly applied in the training stage of the network. When the network training is completed and tracking is performed, the guiding branch will be blocked.

[0126] In the training stage, first, the template image, the search region image, and the padding image are respectively input into the template branch, the search region branch, and the guiding branch to extract features. The images of each branch are data processed by the sampling strategy in advance, and the sizes are fixed as 127×127×3, 255×255×3, and 255×255×3 respectively. Then, depthwise cross-correlation calculation will be performed among the feature tensors output from each branch. Finally, the loss function is calculated by using formulas (1.2) and (1.3), and the network model is guided to complete the training.

[0127] In the tracking stage, the feature extraction unit first inputs the template image and the search area image into the template branch and the search area branch respectively to extract features. Then, the response map acquisition unit performs cross-correlation calculation using the feature tensors output by the two branches to obtain the final response map. Finally, the position localization unit finds the maximum peak in the obtained response map and maps it to the search area image to determine the exact position of the target. To solve the problem of target scale variation, the present invention stipulates that the mapped scale estimates are three kinds: 1.0375{-1,0,1}.

[0128] Embodiment 3

[0129] A storage medium stores program files capable of implementing the foreground information-guided Siamese convolutional neural network target tracking method of any one of the above.

[0130] Embodiment 4

[0131] A processor is used to run a program, wherein when the program runs, it executes the foreground information-guided Siamese convolutional neural network target tracking method of any one of the above.

[0132] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0133] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0134] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the system embodiments described above are only illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0135] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0137] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0138] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A foreground information-guided twin convolutional neural network object tracking method, characterized in that It includes the following steps: Replace the shallow network with an improved residual network to obtain depth features with higher semantic information content; Adopt a sampling strategy to enhance the challenging factors in positive sample pairs; Input the template image, search region image, and padding image into the template branch, search region branch, and guiding branch respectively to extract features, and perform depth cross-correlation calculation among the feature tensors output by each branch. In the calculation, use the padding loss calculation method for foreground information to improve the loss function; Based on the improved loss function, introduce the padding loss and guide the network model to complete training; where: The replacement of the shallow network with the improved residual network includes: For the improved network, change the stride of the last two convolutional blocks in the residual network to 1, and add a convolutional layer with a kernel size of 1×1 at the end of the network to reduce the output dimension; The adoption of the sampling strategy to enhance the challenging factors in positive sample pairs includes: The sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference; The sampling strategy adopts a random affine transformation combining rotation and shear mapping to simulate the deformation of the target; The use of the padding loss calculation method for foreground information in the calculation to improve the loss function includes: Fully utilize the foreground information and use the padding loss calculation method to improve the loss function of the method. The input of the guiding branch is a padding image with a size of 255×255×3; this padding image is obtained by filling the background with the mean value of the whole image, and the filling method is: In the formula, a and b represent the a-th row and b-th column in the image, fg represents foreground information, and bg represents background information; and the network model used for feature extraction in the guiding branch is the same as that in the target template branch and search region branch, and all adopt the improved ResNet-18 network as the feature extraction network; When the template image, search region image, and padding image are given, input each image into the corresponding branch to extract convolutional features; then perform depth cross-correlation calculation among the feature tensors output by each branch; where the depth cross-correlation result between the target template branch and the search region branch is set as S, and the depth cross-correlation result between the target template branch and the guiding branch is S′.

2. The foreground information-guided twin convolutional neural network object tracking method according to claim 1, wherein The method further includes the steps: In the tracking stage, input the template image and search region image into the template branch and search region branch respectively to extract features; Use the feature tensors output by the two branches to perform cross-correlation calculation to obtain the final response map; Find the maximum peak in the obtained response map and map it to the search region image to determine the exact position of the target.

3. The foreground information-guided twin convolutional neural network object tracking method according to claim 1, wherein The sampling strategy uses an artificially designed occlusion mask to increase the number of positive sample pairs under occlusion interference includes: In the offline training stage of the network model, the sampling strategy regards the template image in the positive sample pair as the object to be augmented; the actual size W×H of the target is obtained according to the target information in the template image; after obtaining the target size, an occlusion mask with a pre-designed direction is generated according to this information, and the size of the mask is fixed at W / 2×H / 2; at the same time, each occlusion mask removes the image pixels it covers, so as to simulate the occlusion interference in the tracking process and improve the robustness when dealing with occlusion.

4. The foreground information-guided twin convolutional neural network object tracking method according to claim 3, characterized in that, The sampling strategy adopts a random affine transformation combining rotation and shear mapping to simulate the deformation of the target, including: When a template image is given, the sampling strategy rotates the template image within an angle range of θ = ±30°; at the same time, the template image will be subjected to shear mapping in the X and Y directions, and the mapping angle ranges in the two directions are φ = ±25° and ψ = ±25° respectively; the above two transformations are randomly combined to form the random affine transformation adopted in the sampling strategy.

5. The foreground information-guided twin convolutional neural network object tracking method according to claim 1, wherein The introduction of the padding loss based on the improved loss function includes: After the depth cross-correlation calculation between branches, the padding loss is introduced on the basis of the loss function of the SiamFC method. The calculation method of the padding loss is: In the formula, C, H, and W respectively represent the number of channels, height, and width of the depth cross-correlation result; the loss function after fusing the padding loss is defined as: loss final =(1 - λ)loss training + λloss padding (1.3) In the formula, λ is the weight for fusing the calculations of each loss.

6. A foreground information-guided Siamese convolutional neural network target tracking device for the foreground information-guided Siamese convolutional neural network target tracking method described in claim 1, characterized in that, Including: A network replacement unit for replacing the shallow network with an improved residual network to obtain deeper features with higher semantic information; A factor improvement unit for using the sampling strategy to improve the challenging factors in the positive sample pair; A cross-correlation calculation unit for inputting the template image, the search region image, and the padding image into the template branch, the search region branch, and the guiding branch respectively to extract features, and performing depth cross-correlation calculation between the feature tensors output by each branch, and using the padding loss calculation method for the foreground information to improve the loss function in the calculation; A loss padding unit for introducing the padding loss on the basis of the improved loss function and guiding the network model to complete training.

7. The foreground information-guided twin convolutional neural network object tracking device according to claim 6, characterized in that The device further includes: A feature extraction unit for inputting the template image and the search region image into the template branch and the search region branch respectively to extract features during the tracking stage; A response map acquisition unit for performing cross-correlation calculation using the feature tensors output by the two branches to obtain the final response map; A position localization unit for finding the maximum peak in the obtained response map and mapping it to the search region image to determine the exact position of the target.

Citation Information

Patent Citations

  • Visual multi-target tracking method and device based on deep learning

    CN111161311A

  • Visual target tracking method of full-convolution integral type and regression twin network structure

    CN111179307A