Object Tracking Method Based on Long- and Short-Term Integrated Appearance Update Mechanism
By introducing long-term and short-term integrated appearance update mechanism and Gaussian fuzzy data enhancement technology in target tracking, the tracking accuracy problems of twin networks when target appearance changes and the existence of interfering objects are solved, achieving higher tracking accuracy and success rate.
Patent Information
- Application Number
- CN202210655087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The existing target tracking method based on twin networks is difficult to accurately infer the target position when the target appearance changes or similar interfering objects appear, which affects the tracking accuracy.
The appearance update mechanism based on long and short time integration is adopted, and by building a twin network, long and short time template groups in the long and short time template pool are used, combined with Gaussian fuzzy data enhancement technology, the target appearance is effectively updated and the distinction between interfering objects is achieved.
It improves the accuracy of target tracking and the ability to distinguish similar interfering objects, and is especially suitable for long-term target tracking tasks in complex scenarios, achieving a tracking success rate of 69.1%.
Smart Images

Figure CN114972435B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, image processing, and deep learning, and particularly relates to an object tracking method based on a long-short-term integrated appearance update mechanism. Background Art
[0002] Object tracking technology is an important research direction in computer vision and has broad application prospects and potential economic value in intelligent monitoring, intelligent visual navigation, intelligent transportation, human-computer interaction, etc. The object tracking task requires continuously annotating the position of an object in subsequent frames based on the initial object position marked by a rectangular box in a given video sequence.
[0003] During the tracking process, due to factors such as illumination, occlusion, motion blur, and deformation, the appearance of the object will change greatly. Often, after tracking for a period of time, the object is difficult to recognize compared with the first frame, which brings great challenges to the object tracking task. At the same time, there will be interfering objects in the video sequence that are similar in appearance to the object, making it impossible for the object tracking method to annotate the correct object position.
[0004] Deep learning-based methods are usually implemented based on Siamese networks, such as SiamFC, SINT, etc., and have achieved good results in terms of tracking accuracy and tracking rate in the field of object tracking. However, these methods usually do not introduce an object appearance update mechanism and an anti-interference mechanism, and cannot accurately infer the object position when the object shape changes or similar interfering objects appear, affecting the tracking accuracy. DiMP achieves better performance by carefully designing a loss function, making full use of the information of the tracking background and updating the model weights in the way of the steepest descent, but this method requires updating the network parameters, reducing the running rate of the tracker. STARK designs a confidence branch for template update to capture the time-varying object appearance, but the condition for judging the update is very simple and cannot effectively distinguish similar objects, resulting in poor-quality or even incorrect object appearance being introduced during tracking. Summary of the Invention
[0005] The purpose of the present invention is to provide an object tracking method based on a long-short-term integrated appearance update mechanism to solve the technical problem of providing a long-short-term integrated appearance update mechanism for a Siamese network tracker, which can effectively update the object appearance, improve the ability to distinguish similar interfering objects, and thus improve the accuracy of object tracking.
[0006] To solve the above technical problem, the specific technical solution of the present invention is as follows:
[0007] An object tracking method based on a long-short-term integrated appearance update mechanism includes the following steps:
[0008] Step S1: Perform data preprocessing steps including data augmentation techniques based on Gaussian blur;
[0009] Step S2: Construct a Siamese network and train the network parameters using backpropagation;
[0010] Step S3: Based on the trained Siamese network, perform object tracking in the video using the long - short - term integrated appearance update mechanism: Input the search region and the long - term template group and short - term template group selected from the long - short - term template pool into the network to obtain the long - term result corresponding to the long - term template group and the short - term result corresponding to the short - term template group; Apply the long - short - term integrated appearance update mechanism, analyze the two prediction results to obtain the position of the object in the current frame; Update the long - short - term template pool.
[0011] Further, step S1 specifically includes the following steps:
[0012] Step S101: Randomly select a video sequence from the dataset and randomly select three frames of images containing the same object from the video sequence;
[0013] Step S102: For these three frames of images, respectively crop out two object templates and a search region image centered on the object. The cropped images are square, that is, the length and width of the two template images are equal, and the length and width of the search region image are equal;
[0014] Step S103: Only during training, apply the data augmentation technique based on Gaussian blur to the two object templates.
[0015] Further, the data augmentation technique based on Gaussian blur in step S103 includes the following steps: Select any square object template, divide it into 4 small squares with the same shape, and randomly select 1 - 2 small squares to apply Gaussian blur.
[0016] Further, step S2 includes:
[0017] Step S201: Construct the network structure of the Siamese network, including a backbone network, a global feature enhancement network, a rectangular box prediction branch, and a confidence branch;
[0018] Step S202: Input the two object templates and a search region obtained in step S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the backbone network, the global feature enhancement network, and the rectangular box prediction branch;
[0019] Step S203: Input the two object templates and a search region obtained in step S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the confidence branch while keeping the parameters of the backbone network, the global feature enhancement network, and the rectangular box prediction branch obtained in step S202 fixed.
[0020] Furthermore, in step S201, the first three-layer network structure of ConvNeXt is used as the backbone network; the Transformer encoder is used as the global feature enhancement network; the fully convolutional network is used as the rectangular box prediction branch; and the convolutional network is used as the confidence branch.
[0021] Furthermore, in step S3, the object tracking steps are as follows:
[0022] Step S301: Input the search region and the long-term template group selected from the long-term and short-term template pools into the network to obtain the rectangular box b output by the rectangular box prediction branch L , and the confidence c output by the confidence branch L ; Input the search region and the short-term template group selected from the long-term and short-term template pools into the network to obtain the rectangular box b output by the rectangular box prediction branch S , and the confidence c output by the confidence branch S ; Among them, the long-term template group contains a fixed target initial template feature and a slowly updated and stable dynamic template feature; the short-term template group contains a fixed target initial template feature and a rapidly iteratively updated dynamic template feature;
[0023] Step S302: Apply the long-term and short-term integrated appearance update mechanism, and use b L , b S , c L , c S , and the rectangular box position b of the previous frame old and confidence c old to obtain the final predicted rectangular box b out , and the update status identifier flag;
[0024] Step S303: If the update condition is met, that is, the update status identifier obtained in step S302 is true, that is, when flag = True, update the long-term template group and the short-term template group respectively: when the confidence c corresponding to the long-term template group L is greater than the corresponding confidence threshold τ L , perform template update at the frame interval of T L ; when the confidence c corresponding to the short-term template group S is greater than the corresponding confidence threshold τ S , perform template update at the frame interval of T S .
[0025] Furthermore, in step S302, first judge IoU(b L , b S ) > τ 1 , Calculate bL and b S The area of the intersection of two rectangular frames, Union(b L and b S ) calculates b L and b S The area of the union of two rectangular frames; if the judgment formula is true, it indicates that the tracking state is good, then let b out = b S , flag = True; if the judgment formula is false, it indicates that the tracking state is poor, then calculate the following judgment formula:
[0026] (1 + IoU(b z and b old ))(1 + c z - c old ), z ∈ (L, S)
[0027] where τ 1 is the intersection-over-union threshold, and the intersection-over-union IoU is used to measure the overlap degree between rectangular frames; the subscript z ∈ (L, S) means that the rectangular frames and confidence results of the long-term template group L and the short-term template group S are respectively brought in for comparison, and the rectangular frame corresponding to the template group with the larger value of this formula after comparison is selected as the final target rectangular frame, and the status identifier flag = False is updated.
[0028] Furthermore, in step S303, for the long-term template group, a fixed confidence threshold τ L is used for template update; for the short-term template group, a dynamically decreasing threshold τ S = τ - βn is used for update, where τ is the initial threshold, β is a positive number, and n is the frame interval between the current frame number and the frame number of the previous update.
[0029] The object tracking method based on the long-term and short-term integrated appearance update mechanism of the present invention has the following advantages:
[0030] The object tracking method based on the long-term and short-term integrated appearance update mechanism proposed by the present invention, using the data preprocessing method in step S1, can mine the global semantic information contained in the data and improve the accuracy of object tracking; using the Siamese network object tracker based on the ConvNeXt convolutional network and Transformer proposed in S2, can improve the feature extraction ability of the backbone network and the object rectangular frame prediction ability based on features; using the tracking process in step S3, can judge the tracking state and guide template update through the prediction result difference between the long-term tracker and the short-term tracker, can effectively distinguish interference objects, improve the tracking accuracy, and is especially suitable for long-term object tracking tasks in complex scenarios. The object tracking method based on the long-term and short-term integrated appearance update mechanism proposed by the present invention can achieve a tracking success rate of 69.1% on the LaSOT dataset. Description of the Drawings
[0031] Figure 1 This is the flow chart of the object tracking method based on the long - short - term integrated appearance update mechanism proposed by the present invention;
[0032] Figure 2 This is an example of the target template image and the search area image cropped by the present invention;
[0033] Figure 3 These are two tracking rectangular frames output by the long - short - term template of the present invention;
[0034] Figure 4(a) is the tracking success rate graph of the tracking method proposed by the present invention in the result comparison of the LaSOT dataset;
[0035] Figure 4(b) is the tracking average precision graph of the tracking method proposed by the present invention in the result comparison of the LaSOT dataset. Detailed Embodiment
[0036] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a target tracking method based on the long - short - term integrated appearance update mechanism of the present invention with reference to the accompanying drawings.
[0037] A target tracking method based on the long - short - term integrated appearance update mechanism, as Figure 1 shown, includes the following steps:
[0038] Step S1: Perform data pre - processing steps including data augmentation techniques based on Gaussian blur, including the following steps:
[0039] Step S101: Randomly select a video sequence from the dataset, and randomly select three frames of images containing the same target from the video sequence.
[0040] Step S102: For these three frames of images, respectively crop out two target templates z1 and z2 and a search area centered on the target R represents the real number field, z1, The shape of the cropped image is square, the length and width of the two template images are equal, that is, H z = W z , the length and width of the search area image are equal, that is, H x = W x .
[0041] Select any square target template, divide it into 4 small squares with the same shape, and randomly select 1 - 2 small squares to apply Gaussian blur.
[0042] Step S103: Only during training, apply the data augmentation technique based on Gaussian blur to the two target templates.
[0043] Step S2: Construct a Siamese network and train the network parameters using backpropagation, including the following steps:
[0044] Step S201: Construct the network structure of the Siamese network, including a backbone network, a global feature enhancement network, a rectangular box prediction branch, and a confidence branch.
[0045] Use the first three layers of the ConvNeXt network structure as the backbone network. The network downsamples the input template image and search region image by s times respectively to obtain two template features and search region features C is the number of feature channels after dimension elevation.
[0046] Use the Transformer encoder as the global feature enhancement network to generate more robust features with the similarity relationship between the template and the search region.
[0047] Use a fully convolutional network as the rectangular box prediction branch to output a heatmap with the corner positions of the target rectangular box, and estimate the position b of the target rectangular box.
[0048] Use a convolutional network as the confidence branch to output the predicted confidence c.
[0049] S202: Input the two target templates and one search region obtained in S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the backbone network, the global feature enhancement network, and the rectangular box prediction branch.
[0050] S203: Input the two target templates and one search region obtained in S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the confidence branch while keeping the parameters of the backbone network, the global feature enhancement network, and the rectangular box prediction branch obtained in the fixed step S202 unchanged.
[0051] Step S3: Based on the trained Siamese network, use the long-short-term integrated appearance update mechanism to perform target tracking in the video: Input the search region and two sets of templates selected from the long-short-term template pool into the network to obtain the long-term result corresponding to the long-term template group and the short-term result corresponding to the short-term template group; Apply the long-short-term integrated appearance update mechanism to analyze the two prediction results to obtain the position of the target in the current frame; Update the long-short-term template pool.
[0052] The target tracking steps are as follows:
[0053] Step S301: Input the search region and the long-term template group selected from the long-short-term template pool into the network to obtain the rectangular box b output by the rectangular box prediction branch L and the confidence c output by the confidence branch L; Input the search area and the short-term template group selected from the long- and short-term template pools into the network to obtain the rectangle b output by the rectangle prediction branch S , and the confidence c output by the confidence branch S . Among them, the long-term template group contains a fixed target initial template feature and a dynamic template feature that updates slowly and is relatively stable; the short-term template group contains a fixed target initial template feature and a dynamic template feature that is updated rapidly and iteratively.
[0054] Step S302: Apply the long- and short-term integrated appearance update mechanism, and use b L , b S , c L , c S , and the rectangle position b old and the confidence c old of the previous frame to obtain the final predicted rectangle b out , and the update status identifier flag.
[0055] First, judge IoU(b L , b S ) > τ 1 , where τ 1 is the intersection over union threshold. The intersection over union (IoU) is used to measure the overlap degree between rectangles: Intersection(b L , b S ) calculates the area of the intersection of the two rectangles b L , b S , and Union(b L , b S ) calculates the area of the union of the two rectangles b L , b S .
[0056] If the judgment formula is true, it indicates that the tracking state is good, then let b out = b S , flag = True.
[0057] If the judgment formula is false, it indicates that the tracking state is poor and there may be interfering targets, then calculate
[0058] (1 + IoU(b z , b old ))(1 + c z - c old ), z ∈ (L, S)
[0059] Among them, the subscript z ∈ (L, S) indicates that the rectangular frames and confidence results of the long-term template group L and the short-term template group S are respectively brought in for comparison, and the rectangular frame corresponding to the template group with a larger value of this formula after comparison is selected as the final target rectangular frame, and flag = False is set.
[0060] S303: If the update condition is satisfied, that is, the update status identifier obtained in step S302 is true and flag = True, update the long-term template group and the short-term template group respectively: when the confidence c corresponding to the long-term template group L is greater than the corresponding confidence threshold τ L , perform template update according to the frame interval of T L ; when the confidence c corresponding to the short-term template group S is greater than the corresponding confidence threshold τ S , perform template update according to the frame interval of T S .
[0061] For the long-term template group, a fixed threshold τ L is adopted. When the confidence is higher than the threshold, that is, c L > τ L , perform long-term template update.
[0062] For the short-term template group, a dynamic threshold τ S = τ - βn that decreases with time is adopted, where τ is the initial threshold, β is a positive number, and n is the frame interval between the current frame number and the frame number of the previous update. When the confidence is higher than the truncated threshold, that is, c S > max(τ S , τ MAX ), perform short-term template update. Among them, τ MAX is a fixed value that constrains the minimum threshold.
[0063] The first embodiment of the present invention:
[0064] An object tracking method based on a long-term and short-term integrated appearance update mechanism, as Figure 1 shown, includes the following steps:
[0065] Step S1: A data preprocessing step including a data augmentation technique based on Gaussian blur.
[0066] Referring to Figure 2 the data image example shown, step S1 includes the following steps:
[0067] Step S101: Randomly select a video sequence from the dataset and randomly select three frames of images containing the same object from it. Specifically, the training dataset consists of the LaSOT, GOT-10k, TrackingNet, and COCO datasets.
[0068] Step S102: For these three frames of images, respectively crop out two target templates \(z_1, z_2\in R\) 128×128×3 and a search area \(x\in R\) centered on the target 320×320×3 . The shape of the cropped image is square.
[0069] Step S103: Only during training, apply data augmentation technology based on Gaussian blur to the two target templates.
[0070] Select any square target template, draw a cross line through the center of the square to divide it into 4 small squares with the same shape, and randomly select 1 - 2 small squares to apply Gaussian blur.
[0071] Apply Gaussian blur to the upper - left, upper - right, lower - left, and lower - right regions with a probability of 0.1, 0.1, 0.1, 0.1, and to the upper - left - upper - right, lower - left - lower - right regions with a probability of 0.15, 0.15.
[0072] The size of the Gaussian blur template is \(5\times5\), and \(\sigma\) is a random number sampled from a uniform distribution of 1 - 10.
[0073] Step S2: Construct a Siamese network and use backpropagation to train the network parameters.
[0074] Step S2 includes the following steps:
[0075] S201: Construct the network structure of the Siamese network, including a backbone network, a global feature enhancement network, a rectangular box prediction branch, and a confidence branch.
[0076] Use the first three - layer network structure of ConvNeXt as the backbone network.
[0077] The network downsamples the input template image and search area image by 16 times respectively, and the number of output channels is 512 finally, obtaining two template features \(f\) z1 , \(f\) z2 \(\in R\) 8×8×512 and search area feature \(f\) x \(\in R\) 20×20×512 .
[0078] Use the Transformer encoder as the global feature enhancement network to generate more robust features with the similarity relationship between the template and the search area.
[0079] Combine the two template features and the search area feature obtained by the backbone network, serialize them along the spatial direction, and combine them into a feature \(F\in R\) (2×8×8+20×20)×512 The feature \(F\) is the feature input into the encoder.
[0080] The encoder contains 6 encoder layers, each layer contains a multi-head self-attention (MHSA) module sub-layer and a feed-forward network (FFN) sub-layer, and layer normalization (LN) and residual connections are used between sub-layers.
[0081] In this embodiment, post-layer normalization is adopted, denoted as F mid as the intermediate variable of the layer, MHSA is the multi-head self-attention operation, FFN is the feed-forward network operation, LN is the layer normalization operation, and the input F between each layer in and the output F out have the following relationship:
[0082] F mid = LN(MHSA(F in ) + F in )
[0083] F out = LN(FFN(F mid ) + F mid )
[0084] Use a fully convolutional network as the rectangular box prediction branch, and output a heat map with the corner positions of the target rectangular box to estimate the position b of the target rectangular box.
[0085] Select the enhanced feature F ∈ R (2×8×8+20×20)×512 output by the Transformer encoder, and for the part belonging to the search area F x ∈ R (20×20)×512 , restore it to a feature F x_reshape ∈ R 20 ×20×512 with the same size as the search area feature output by the backbone network, and input it into the fully convolutional network.
[0086] The fully convolutional network contains two branches, which respectively predict the upper left corner point and the lower right corner point of the rectangular box.
[0087] Each branch contains 5 convolutional layers, the convolutional kernel size of each convolutional layer is 3×3, the stride is 1, and the number of channels is 256, 128, 64, 32, 1 respectively.
[0088] The heat map finally output by the branch passes through the sigmoid activation function to obtain the probability distribution P(x, y) of the corner points, and the finally predicted corner point coordinates are the probability weighting on the entire heat map of size H×W, and the formula is as follows:
[0089]
[0090]
[0091] Use a convolutional network as the confidence branch to output the predicted confidence c.
[0092] Use the same input feature F as the rectangular box prediction branch in step S201 x_reshape ∈R 20×20×512 , passing through two convolutional layers, each with a convolutional kernel size of 1×1, a stride of 1, and the number of channels being 256 and 128 respectively. Then pass through a global average pooling layer, and then through two convolutional layers, each with a convolutional kernel size of 1×1, a stride of 1, and the number of channels being 64 and 1 respectively. After passing through the sigmoid activation function, the confidence score of this prediction in the 0-1 interval is obtained.
[0093] Step S202: Input the two target templates and a search area obtained in step S1 into the siamese network, and use the AdamW adaptive learning rate gradient descent method to train the backbone network, the global feature enhancement network, and the rectangular box prediction branch.
[0094] Use Python 3.6 and Pytorch 1.7 as the training language and deep learning tool, and use 4-GPU distributed training.
[0095] Train for a total of 600 epochs, with 60,000 samples in each epoch, and the batch size on each GPU is 8.
[0096] Adopt the AdamW optimizer, with an initial learning rate of 0.00004 and a weight decay of 0.0001. The learning rate will be reduced by a factor of ten after 400 epochs. Adopt the GIOU loss function.
[0097] Step S203: Input the two target templates and a search area obtained in step S1 into the siamese network, and use the AdamW adaptive learning rate gradient descent method to train the confidence branch while keeping the parameters of the backbone network, the global feature enhancement network, and the rectangular box prediction branch network fixed.
[0098] Use Python 3.6 and Pytorch 1.7 as the training language and deep learning tool, and use 4-GPU distributed training.
[0099] Train for a total of 40 epochs, with 60,000 samples in each epoch, and the batch size on each GPU is 8.
[0100] Adopt the AdamW optimizer, with an initial learning rate of 0.00004 and a weight decay of 0.0001. Freeze the parameters of the backbone network, the global feature enhancement network, and the rectangular box prediction branch network during training. Adopt the binary classification loss function.
[0101] Step S3: Based on the trained Siamese network, use the long-short-term integrated appearance update mechanism to perform object tracking in the video: input the search region and the long-term template group and short-term template group in the long-short-term template pool into the network to obtain the long-term result corresponding to the long-term template group and the short-term result corresponding to the short-term template group; apply the long-short-term integrated appearance update mechanism, analyze the two prediction results, obtain the position of the object in the current frame; update the long-short-term template pool.
[0102] In step S3, the object tracking steps are as follows:
[0103] Step S301: Input the search region and the long-term template group selected from the long-short-term template pool into the network to obtain the rectangle b output by the rectangle prediction branch L , and the confidence c output by the confidence branch L ; input the search region and the short-term template group selected from the long-short-term template pool into the network to obtain the rectangle b output by the rectangle prediction branch S , and the confidence c output by the confidence branch S .
[0104] The long-term template group contains a fixed target initial template feature and a dynamically updated and relatively stable dynamic template feature, which reduces the noise introduced in template updates.
[0105] The template image is obtained by the cropping method in step S102, and the template image is converted into a template feature by the backbone network in step S201. The target initial template feature is obtained from the template image of the first frame, and the dynamic template feature is obtained by the long-term template group update method in step S303.
[0106] The short-term template group contains a fixed target initial template feature and a dynamically updated and relatively stable dynamic template feature, which is used to record the rapid changes of the target.
[0107] The template image is obtained by the cropping method in step S102, and the template image is converted into a template feature by the backbone network in step S201. The target initial template feature is obtained from the template image of the first frame, and the dynamic template feature is obtained by the short-term template group update method in step S303.
[0108] Step S302: Apply the long-short-term integrated appearance update mechanism, use b L , b S , c L , c S , and the rectangle position b old and confidence c old of the previous frame to obtain the final predicted rectangle b out , and the update status identifier flag.
[0109] First, determine whether IoU(b L , b S ) > τ 1 , where τ 1 = 0.5. The intersection over union (IoU) is used to measure the overlapping degree between rectangular boxes: Intersection(b L , b S ) calculates the area of the intersection of the two rectangular boxes b L , b S , and Union(b L , b S ) calculates the area of the union of the two rectangular boxes b L , b S .
[0110] If the judgment formula is true, it indicates that the tracking state is good, then let b out = b S , and flag = True.
[0111] If the judgment formula is false, it indicates that the tracking state is poor, and there may be interfering targets. Then calculate (1 + IoU(b z , b old ))(1 + c z - c old ), where z ∈ (L, S). Here, the subscript z ∈ (L, S) means that the rectangular boxes and confidence results of the long-term template group L and the short-term template group S are respectively brought in for comparison, and the rectangular box corresponding to the template group with the larger value after comparison is selected as the final rectangular box, and let flag = False.
[0112] Step S303: If the update condition is met, that is, the update status identifier obtained in step S302 is true, flag = True, update the long-term template group and the short-term template group respectively: When the confidence c L corresponding to the long-term template group is greater than the corresponding confidence threshold τ L , perform template update at the frame interval of T L ; when the confidence c S corresponding to the short-term template group is greater than the corresponding confidence threshold τ S , perform template update at the frame interval of T S .
[0113] For the long-term template group, a fixed threshold τ L = 0.5 is adopted. When the confidence is higher than the threshold, that is, c L > τ L , perform long-term template update. The frame interval T L takes the value of 400.
[0114] For the short-term template group, a dynamic threshold τ that decreases over time is adopted S = τ - βn, where τ = 0.95, β = 0.005, and n is the frame interval between the current frame number and the frame number of the previous update. When the confidence level is higher than the threshold, i.e., c S > max(τ S , τ MAX ), short-term template update is performed. Among them, τ MAX is a fixed value that constrains the minimum threshold. Specifically, τ MAX = 0.5, and the frame interval T S takes a value of 400.
[0115] As Figure 2 shown are examples of the target template image and the search area image cropped by the present invention. As Figure 3 shown are two tracking rectangular frames output by the long-term and short-term templates of the present invention.
[0116] As shown in Fig. 4(a) is the tracking success rate graph of the tracking method proposed by the present invention in the result comparison of the LaSOT dataset. As shown in Fig. 4(b) is the tracking average precision graph of the tracking method proposed by the present invention in the result comparison of the LaSOT dataset. The target tracking method based on the long-term and short-term integrated appearance update mechanism proposed by the present invention can utilize the data preprocessing method in step S1 to mine the global semantic information contained in the data and improve the accuracy of target tracking; by using the Siamese network target tracker based on the ConvNeXt convolutional network and Transformer in S2, it can enhance the feature extraction ability of the backbone network and the ability to predict the target rectangular box based on features; by using the tracking process in step S3, it can judge the tracking state and guide template update based on the prediction result difference between the long-term tracker and the short-term tracker, effectively distinguish interfering targets, and improve the tracking accuracy, especially suitable for long-term target tracking tasks in complex scenarios. The target tracking method based on the long-term and short-term integrated appearance update mechanism proposed by the present invention can achieve a tracking success rate of 69.1% on the LaSOT dataset.
[0117] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A target tracking method based on a long - short - term integrated appearance update mechanism, characterized in that, it includes the following steps: Step S1: Perform data pre - processing steps including data augmentation techniques based on Gaussian blur; Step S2: Construct a Siamese network and train network parameters using backpropagation; Step S3: Based on the trained Siamese network, use the long - short - term integrated appearance update mechanism to perform target tracking in the video: Input the search region and the long - term template group and short - term template group selected from the long - short - term template pool into the network to obtain the long - term result corresponding to the long - term template group and the short - term result corresponding to the short - term template group; Apply the long - short - term integrated appearance update mechanism to analyze the two prediction results to obtain the position of the target in the current frame; Update the long - short - term template pool; In the said Step S3, the target tracking steps are as follows: Step S301: Input the search region and the long-term template group selected from the long-term and short-term template pools into the network to obtain the rectangle b output by the rectangle prediction branch L , and the confidence c output by the confidence branch L ; Input the search region and the short-term template group selected from the long-term and short-term template pools into the network to obtain the rectangle b output by the rectangle prediction branch S , and the confidence c output by the confidence branch S ; Among them, the long-term template group includes a fixed target initial template feature and a slowly updated and stable dynamic template feature; the short-term template group includes a fixed target initial template feature and a rapidly iteratively updated dynamic template feature; Step S302: Apply the long - short - term integrated appearance update mechanism, and use b L , b S , c L , c S , and the position b old of the rectangular box in the previous frame and the confidence c old to obtain the final predicted rectangular box b out , and the update status identifier flag; Step S303: If the update condition is satisfied, that is, the update status identifier obtained in step S302 is true, i.e., flag = True, update the long-term template group and the short-term template group respectively: When the confidence level c of the long-term template group L is greater than the corresponding confidence level threshold τ L at this time, perform template update according to the frame interval of T L ; When the confidence level c of the short-term template group S is greater than the corresponding confidence level threshold τ S at this time, perform template update according to the frame interval of T S ; In step S302, first, it is determined whether IoU(b L ,b S )>τ 1 . Intersection(b L ,b S ) calculates the area of the intersection of the two rectangular frames b L ,b S , and Union(b L ,b S ) calculates the area of the union of the two rectangular frames b L ,b S . If the judgment formula is true, it indicates that the tracking state is good, and then b out = b S , and flag = True; if the judgment formula is false, it indicates that the tracking state is poor, and then the following judgment formula is calculated: (1 + Iou(b z , b old ))(1 + c z -c old ), z ∈ (L, S) where τ 1 is the intersection over union (IoU) threshold, and the IoU is used to measure the overlapping degree between rectangular boxes; the subscript z ∈ (L, S) indicates that the rectangular boxes and confidence results of the long-term template group L and the short-term template group S are respectively brought in for comparison, and the rectangular box corresponding to the template group with a larger value of this formula after comparison is selected as the final target rectangular box, and the status identifier flag = False is updated; In step S303, for the long-term template group, a fixed confidence threshold τ is adopted L to update the template; for the short-term template group, a dynamically decreasing threshold τ S = τ - βn is adopted for update, where τ is the initial threshold, β is a positive number, and n is the frame interval between the current frame number and the frame number at the previous update.
2. The target tracking method based on the long - short - term integrated appearance update mechanism according to claim 1, characterized in that, the said Step S1 specifically includes the following steps: Step S101: Randomly select a video sequence from the dataset and randomly select three frames of images containing the same target from the video sequence; Step S102: For these three frames of images, respectively crop out two target templates and a search region image centered on the target. The cropped images are in the shape of a square, that is, the lengths and widths of the two template images are equal, and the lengths and widths of the search region images are equal; Step S103: Only during training, apply the data augmentation technique based on Gaussian blur to the two target templates.
3. The target tracking method based on the long - short - term integrated appearance update mechanism according to claim 2, characterized in that, the data augmentation technique based on Gaussian blur in the said Step S103 includes the following steps: Select any square target template, divide it into 4 small squares with the same shape, and randomly select 1 - 2 small squares to apply Gaussian blur.
4. The target tracking method based on the long - short - term integrated appearance update mechanism according to claim 3, characterized in that, the said Step S2 includes: Step S201: Construct the network structure of the Siamese network, including a backbone network, a global feature enhancement network, a rectangular box prediction branch, and a confidence branch; Step S202: Input the two target templates and a search region obtained in Step S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the backbone network, the global feature enhancement network, and the rectangular box prediction branch; Step S203: Input the two target templates and a search region obtained in Step S1 into the Siamese network, and use the AdamW adaptive learning rate gradient descent method to train the confidence branch while keeping the parameters of the backbone network, the global feature enhancement network, and the rectangular box prediction branch obtained in Step S202 fixed.
5. The target tracking method based on the long - short - term integrated appearance update mechanism according to claim 4, characterized in that, in the said Step S201, use the first three - layer network structure of ConvNeXt as the backbone network; use the Transformer encoder as the global feature enhancement network; use the fully convolutional network as the rectangular box prediction branch; use the convolutional network as the confidence branch.
Citation Information
Patent Citations
Template update target tracking algorithm based on multilayer features of fully convolutional twin network
CN114581486A
Spiking neural network-based short-range tracking method and system
WO2021012752A1