A video object tracking method based on a promptable segmentation model

By adopting a method based on a promptable segmentation model in the video target tracking technology, an object-level feature extraction and self-optimization target box decoder is realized, solving the problem of background noise interference and lack of self-optimization of bounding box prediction heads in the prior art, and improving the robustness and accuracy of tracking.

CN117173219BActive Publication Date: 2025-05-27ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311240754.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2025-05-27
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

In the existing video target tracking technology, the interaction between the template frame and the search frame is image-level, resulting in background noise interference, and the bounding box prediction head lacks self-optimization capabilities, resulting in insufficient tracking accuracy and robustness.

Method used

Using a video target tracking method based on a promptable segmentation model, an object-level feature extraction and self-optimization target box decoder is realized by constructing a video single-object tracking encoder-feature enhancement-decoder paradigm, and a self-attention enhancement unit of templates and search area features and a target-oriented foreground prompting unit.

Benefits of technology

It improves the robustness and accuracy of video target tracking, reduces the impact of background information in template images on tracking, and achieves efficient tracking in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173219B_ABST
    Figure CN117173219B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of video object tracking, and proposes a video object tracking method based on a promptable segmentation model, including the following steps: S1. Construct a video single-object tracking encoder-feature enhancement-decoder paradigm; S2. Construct an encoder based on a promptable segmentation model; S3. Construct a self-attention enhancement unit for template and search region features; S4. Construct a target-oriented foreground prompt unit; S5. Construct a self-optimizable target box decoder; S6. After constructing a single-object tracking model including S1-S5, train the single-object tracking model on a server, and optimize the network parameters by reducing the overall loss value of the network loss function until the network converges; S7. Use the trained network model to track a specified single object in the video sequence to be tracked. The present invention can robustly and efficiently achieve video single-object tracking in complex scenarios, and achieves better object tracking effects compared with other methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video target tracking, and in particular to a video target tracking method based on a promptable segmentation model. Background Art

[0002] Video target tracking is one of the hot issues in the field of computer vision because it has been widely studied and applied in many industries and fields, such as intelligent video surveillance, autonomous driving, etc. The video target tracking task aims to track the target in the video through the first frame image of the video and its initial bounding box of the specified target. The main technical difficulties of the video target tracking task lie in the continuous change and arbitrariness of the target, the occlusion of the tracked target by other objects, and the rapid movement of the tracked target. This causes the appearance of the tracked target to change greatly in each frame and is easily affected by the appearance of the surrounding environment.

[0003] An existing video object tracking paradigm is Figure 1 As shown in a, the twin network is first used as an image encoder to extract features from the search frame and the template frame, then the extracted features are interacted with the features of the search frame and the template frame, and finally the interactive features are sent to the bounding box prediction head to obtain the prediction results. This paradigm has two major problems:

[0004] (1) The interaction between the template frame and the search frame is at the image level rather than the object level, which inevitably introduces some background noise in the template frame, causing the model to mistakenly believe that this part of the background noise is also the target to be tracked. For video tracking tasks, because each subsequent frame will be compared with the template image, the information contained in the template image is crucial. Therefore, under this paradigm, the detailed background information in the template image will be mistakenly considered to be an indispensable part of the tracked target, which can easily lead to incorrect tracking during tracking.

[0005] (2) The bounding box prediction head does not have the ability to self-optimize. This tracking paradigm directly obtains the bounding box through the prediction head and adjusts the parameters of the prediction head through neural network optimization. This method does not allow the prediction head to understand its own input-output relationship, and thus cannot understand the quality of its output bounding box and how to adjust its own parameters.

[0006] There is also a video object tracking paradigm such as Figure 1As shown in Figure b, using the Vision Transformer (ViT), the search frame and the template frame are first mapped into small image patches. After encoding the image patches, the search frame image patches and the template frame image patches are concatenated together and then fed into a series of Vision Transformer encoding blocks for feature extraction and feature interaction. The disadvantages of this paradigm also include that the interaction between the template frame and the search frame is at the image level rather than the object level, and the bounding box prediction head does not have the ability of self-optimization. Summary of the Invention

[0007] In view of the above problems, the present invention provides a video object tracking method based on a promptable segmentation model to solve the problems faced by the above video object tracking paradigms. This method can ensure high tracking accuracy and speed in the actual scenarios of complex environments, and track the selected target intelligently, quickly and accurately.

[0008] To achieve the above object, the present invention provides a video object tracking method based on a promptable segmentation model, including the following steps:

[0009] S1. Construct a video single-object tracking encoder-feature enhancement-decoder paradigm;

[0010] S2. Construct an encoder based on a promptable segmentation model;

[0011] S3. Construct a self-attention enhancement unit for template and search region features;

[0012] S4. Construct a target-oriented foreground prompt unit;

[0013] S5. Construct a self-optimizable target box decoder;

[0014] S6. Under the video single-object tracking encoder-feature enhancement-decoder paradigm, construct a single-object tracking model including an encoder based on a promptable segmentation model, a self-attention enhancement unit for template and search region features, a target-oriented foreground prompt unit, and a self-optimizable target box decoder, and train the single-object tracking model on the server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges;

[0015] S7. Use the trained network model to track a specified single target in the video sequence to be tracked.

[0016] Preferably, the step S1 specifically includes the following steps:

[0017] S11. Based on the image encoder of the promptable segmentation model, establish an image feature extractor by combining an adapter, a commonly used tool in the field of natural language processing, and input the search frame and the template frame Extract features using the image feature extractor to obtain the search frame feature x s and the template frame feature x t ;

[0018] S12. Use the self-attention enhancement unit for the template and search region features to efficiently fuse the search frame feature x s and the template frame feature x t to obtain the enhanced search frame feature F s and the enhanced template frame feature F t ;

[0019] S13. In the target-oriented foreground prompting unit, use the enhanced template frame feature F t , perform segmentation through the promptable segmentation model to obtain the segmentation mask of the object to be tracked in the video object tracking task, and use the segmentation mask to obtain the feature f of the target object obj ;

[0020] S14. Input the feature f of the target object obj and the enhanced search frame feature F s into the self-optimizable target box decoder to obtain the tracking result, where the tracking result includes the target classification score map P, the local offset map O, and the normalized size map S of the tracked target

[0021] Preferably, step S2 specifically includes the following steps:

[0022] S21. Retrain the absolute position encoding and relative position encoding of the vision Transformer in the image encoder of the promptable segmentation model to match the image resolutions of 256×256 and 384×384;

[0023] S22. Use the adapter to cooperate with the image encoder in the promptable segmentation model as the image feature extractor, and use the image feature extractor to extract features from the input search frame and the template frame respectively to obtain the search frame feature and the template frame feature That is:

[0024]

[0025]

[0026] Among them, the role of PatchEmbed(·) is to map image patches to the hidden space is the absolute position encoding, LN(·) is the layer normalization, MSA(·) is the multi-head attention mechanism, Adapter(·) is the adapter, and num_block represents the number of unit blocks in the image feature extractor.

[0027] Preferably, the adapter consists of a fully connected downsampling layer, an activation layer, a fully connected upsampling layer, and a residual connection.

[0028] Preferably, the self-attention enhancement unit of the template and search region features consists of 2 self-attention units, and each self-attention unit consists of a multi-head attention layer (MSA), layer normalization (LN), a multi-layer perceptron (MLP), and a residual connection.

[0029] Preferably, step S3 specifically includes the following steps:

[0030] S31. Concatenate the search frame features obtained in step S22 and the template frame features and add the absolute position encoding Then send it to 2 self-attention units for processing;

[0031] S32. Send the concatenated features to the self-attention enhancement unit of the template and search region features, and split the obtained result to obtain the enhanced search frame features and the enhanced template frame features

[0032] Preferably, the specific implementation process of step S4 is as follows: Use the target box of the template frame as a prompt, and extract features using the prompt encoder in the promptable segmentation model According to the features prompted by the target box and the enhanced template frame features Use the mask decoder in the promptable segmentation model to obtain the segmentation mask m t of the target object, then use the screening ability of the mask to act on the enhanced template frame features, and obtain the target object feature vector through average pooling

[0033] Preferably, step S5 specifically includes the following steps:

[0034] S51. Calculate the cosine similarity between the target object feature vector and the enhanced search frame features to obtain the similarity map between the target object and the search image

[0035] S = f obj ·F s

[0036] Each element S in the similarity graph S (C,i,j) (0 < i < H, 0 < j < W) represents the similarity between the target object feature vector and each image patch of the search frame. S (C,i,j) The larger the value, the higher the similarity, that is, this position can be considered as the foreground point in the video object tracking task; conversely, S (C,i,j) The smaller the value, the lower the similarity, that is, this position can be considered as the background point in the video object tracking task. Select S (C,i,j) The points corresponding to the maximum value on the image and the points corresponding to the minimum value on the image as a pair of positive and negative position point pairs, denoted as P maxmin =(P max , P min );

[0037] S52. Use the positive and negative position point pair P maxmin as a hint to decode the enhanced search frame feature and output the target-oriented search frame feature That is:

[0038]

[0039] S53. Use the target-oriented search frame feature to obtain the target classification score map local offset map normalized size map of the target through a fully convolutional network. The center point of the target object corresponds to the highest score in the classification map. Based on the target classification score map, local offset map, and normalized size map, the preliminary target object bounding box bbox 0 can be obtained;

[0040] S54. Use the preliminary target object bounding box bbox 0 as a hint to input the target-oriented search frame feature into the mask decoder of the hintable segmentation model and output the final search frame feature

[0041] S55. Use the final search frame feature to obtain the target classification score map local offset map normalized size map Based on the target classification score map, local offset map, and normalized size map, the final target object bounding box bbox is obtained, achieving the effect of self-optimization.

[0042] Preferably, step S6 specifically includes the following steps:

[0043] S61. Use the server to perform image cropping and data augmentation: The cropping method is as follows: Crop a rectangular image centered on the target area. The length and width of this rectangular image are 2 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the boundary of the original video is filled with the pixel average value. Finally, scale this rectangular image to 256×256 or 384×384 to form a template frame image; Crop a rectangular image centered on the target area. The length and width of this rectangular image are 4 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the boundary of the original video is filled with the pixel average value. Finally, scale this rectangular image to 256×256 or 384×384 to form a search frame image. Perform data augmentation on the template frame and the search frame, including operations such as random inversion, random grayscale conversion, center point position jitter, and bounding box size jitter;

[0044] S62. Use the server to execute step S2, that is, based on the encoder of the promptable segmentation model, perform feature extraction on both the search frame image and the template frame image; then execute step S3, and perform feature interaction and enhancement on the template frame image feature and the search frame image feature through the self-attention enhancement unit of the template and search area features; then execute step S4 to obtain the target-oriented foreground prompt; finally execute step S5, and obtain the target classification score map, local offset map, and normalized size map of the tracked target through a self-optimizable target box decoder, and then the target object bounding box can be obtained;

[0045] S63. Use the server to train the network model, and perform training in an end-to-end manner. The overall expression of the loss function is:

[0046]

[0047] Among them, L cls represents the focal loss for classification; L iou and L 1 are both used to supervise the bounding box regression, representing the generalized intersection over union loss and the L1 loss respectively. λ iou and represent the weights of the generalized intersection over union loss and the L1 loss respectively, and are set to: λ iou = 2,

[0048] S64. Train the single-object tracking model on the server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges, and obtain the locally optimal network parameters.

[0049] Preferably, step S7 specifically includes the following steps:

[0050] S71. For a given video, initialize the first frame of the video as the template frame for video object tracking, and give the bounding box of the template frame. According to this bounding box, divide a region of the template frame with a preset size change as the template frame image, and the video object tracking network starts tracking from the second frame;

[0051] S72. For each frame from the second frame and later, use the center point of the target box of the previous frame as a reference, and divide a region with a preset size and distance as the search frame image. Input the template frame image and the search frame image into the network model, and output the tracking result of the current frame.

[0052] Compared with the prior art, the beneficial effects of the present invention are:

[0053] A video object tracking method based on a promptable segmentation model provided by the present invention obtains object-level features of the tracking target by using the promptable segmentation model, expands the interaction range of the search frame from the image-level template frame to the object-level target object, and the tracking target features provide important prompt information for the search frame in the target box decoding stage, making the search frame pay more attention to the target object to be tracked and being immune to the background information in the template, thereby improving the robustness of the tracking. In addition, the target box decoder has the ability of self-optimization, can timely feedback and optimize the output bounding box, and improves the accuracy of the tracking. The present invention can accurately and stably track the target in many difficult actual scenarios, and achieves better target tracking effects compared with other methods. Description of the Drawings

[0054] Figure 1 a is a diagram of a classic two-stream video object tracking paradigm;

[0055] Figure 1 b is a diagram of a classic single-stream video object tracking paradigm;

[0056] Figure 2 is the overall structure diagram of a video object tracking method based on a promptable segmentation model of the present invention;

[0057] Figure 3 a is the structure diagram of the adapter used in the present invention;

[0058] Figure 3 b is the structure diagram of a classic Vision Transformer (ViT) block;

[0059] Figure 3 c is the structure diagram of the Vision Transformer (ViT) block using the adapter of the present invention;

[0060] Figure 4 Structural diagram of the target-oriented foreground prompting unit of the present invention;

[0061] Figure 5 Structural diagram of the self-optimizable target box decoder of the present invention. Detailed implementation manners

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] The promptable segmentation model is a large-scale basic model for image segmentation. It supports flexible prompt information and can segment images in real time, with powerful image segmentation capabilities. The promptable segmentation model is trained using 11 million images and 110 million image masks, can generate high-quality masks, and can achieve zero-shot segmentation in general scenarios. For the two video object tracking paradigms in the prior art, the common difficulty is that the interaction between the search frame and the template frame is at the image level rather than the object level. The powerful zero-shot segmentation ability of the promptable segmentation model can provide the object-level mask of the tracking target in video object tracking. Therefore, the promptable segmentation model has the potential to be applied to video object tracking tasks.

[0064] In view of the problems and deficiencies in the prior art, the present invention proposes a video object tracking method based on a promptable segmentation model, which mainly includes the steps of implementing seven stages: the design of the video single-object tracking encoder-feature enhancement-decoder paradigm, the design of an encoder based on a promptable segmentation model, the design of a self-attention enhancement unit for template and search region features, the design of a target-oriented foreground prompting unit, the design of a self-optimizable target box decoder, model training, and model inference.

[0065] A video object tracking method based on a promptable segmentation model proposed by the present invention, as Figure 2 shown, includes the following steps:

[0066] S1. Construct a video single-object tracking encoder-feature enhancement-decoder paradigm;

[0067] S2. Construct an encoder based on a promptable segmentation model;

[0068] S3. Construct a self-attention enhancement unit for template and search region features;

[0069] S4. Construct a target-oriented foreground prompting unit;

[0070] S5. Construct a self-optimizing target box decoder;

[0071] S6. Under the video single-object tracking encoder-feature enhancement-decoder paradigm, construct a single-object tracking model that includes an encoder based on a promptable segmentation model, a self-attention enhancement unit for template and search region features, a target-oriented foreground prompt unit, and a self-optimizing target box decoder, and train the single-object tracking model on a server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges;

[0072] S7. Use the trained network model to track a specified single object in the video sequence to be tracked.

[0073] The following is a detailed description of each step.

[0074] Step S1. Construct a video single-object tracking encoder-feature enhancement-decoder paradigm.

[0075] By changing the existing two-stream model paradigm, add an adapter to the image encoder, and perform feature enhancement such as feature fusion and prompt generation on the features obtained by the image encoder; add a self-optimizing target box feedback method to the decoder.

[0076] Specifically, it mainly includes the following steps:

[0077] S11. Based on the image encoder of the promptable segmentation model, establish an image feature extractor by combining the commonly used tool - Adapter in the field of natural language processing, and input the search frame and the template frame Extract features with this image feature extractor to obtain the search frame feature x s and the template frame feature x t ;

[0078] S12. Use the self-attention enhancement unit for template and search region features to efficiently fuse the search frame feature x s and the template frame feature x t to obtain the enhanced search frame feature F s and the enhanced template frame feature F t ;

[0079] S13. In the target-oriented foreground prompt unit, use the enhanced template frame feature F t to perform segmentation through the promptable segmentation model to obtain the segmentation mask of the object to be tracked in the video object tracking task, and use the segmentation mask to obtain the feature f obj of the target object;

[0080] S14. Input the feature f of the target object obj and the enhanced search frame feature F s into a self-optimizable target box decoder to obtain a tracking result, which includes the target classification score map P, the local offset map O, and the normalized size map S of the tracked target.

[0081] Step S2. Construct an encoder based on a promptable segmentation model. Extract features from the search frame image and the template frame image.

[0082] Specifically, it mainly includes the following steps:

[0083] S21. Since the promptable segmentation model is trained based on images with a resolution of 1024×1024, and the present invention is trained based on images with resolutions of 256×256 and 384×384, re-train the absolute position encoding and relative position encoding of the vision Transformer in the image encoder of the promptable segmentation model to match the image resolutions of 256×256 and 384×384;

[0084] S22. To better transfer the promptable segmentation model to the video object tracking task, the present invention adds an adapter in the image encoder of the promptable segmentation model, thereby constructing the image feature extractor in the present invention. As shown in Figure 3 a, the adapter consists of a fully connected downsampling layer, an activation layer, a fully connected upsampling layer, and a residual connection. During network training, freeze the image encoder in the image feature extractor, and only the adapter needs to be trained. The structure of a classic vision Transformer (ViT) block is shown in Figure 3 b; To achieve better task transfer, the structure of the vision Transformer (ViT) block with an added adapter is shown in Figure 3 c. Use the adapter in combination with the image encoder in the promptable segmentation model as the image feature extractor, and input the search frame and the template frame respectively extract features with this image feature extractor to obtain the search frame feature and the template frame feature That is:

[0085]

[0086]

[0087] Among them, the function of PatchEmbed(·) is to map image patches to the hidden space, is the absolute position encoding, LN(·) is layer normalization, MSA(·) is the multi-head attention mechanism, Adapter(·) is the adapter, and num_block represents the number of unit blocks in the image feature extractor.

[0088] Step S3: Construct a self-attention enhancement unit for the template and search region features.

[0089] The self-attention enhancement unit for the template and search region features consists of 2 self-attention units, and each self-attention unit is composed of a multi-head attention layer (MSA), layer normalization (LN), a multi-layer perceptron (MLP), and a residual connection.

[0090] Specifically, it mainly includes the following steps:

[0091] S31: Take the search frame features and the template frame features obtained in step S22, concatenate them and add the absolute position encoding,

[0092]

[0093]

[0094]

[0095] S32: Feed the concatenated features into the self-attention enhancement unit for the template and search region features, and split the resulting result to obtain the enhanced search frame features and the enhanced template frame features

[0096]

[0097] where split(·) splits the enhanced concatenated features according to the tensor shapes of the search frame features and template frame features before enhancement.

[0098] In this embodiment, the self-attention enhancement unit for the template and search region features in step S3 adds a skip connection to accelerate the convergence of the network.

[0099] Step S4: Construct a target-guided foreground prompting unit.

[0100] In the video object tracking task, the target box of the template frame is known. Therefore, using the powerful inference ability of the promptable segmentation model, the target box of the template frame is used as a prompt.

[0101] As shown Figure 4 below, the specific implementation process is as follows: Use the target box of the template frame as a prompt, and extract features using the prompt encoder in the promptable segmentation model That is

[0102] f bbox = PromptEncoder(bbox)

[0103] According to the features prompted by the target box and the enhanced template frame features Use the mask decoder in the promptable segmentation model to obtain the segmentation mask m of the target object t , that is:

[0104] m t = MaskDecoder(F t , f bbox )

[0105] Then, use the screening ability of the mask to act on the enhanced template frame features, and obtain the target object feature vector through average pooling

[0106] f obj = Avg(select(F t , m t ))

[0107] This feature is the object-level feature, which is convenient for providing more accurate prompt information for the decoder

[0108] Step S5, construct a self-optimizing target box decoder

[0109] As Figure 5 shown, specifically, it mainly includes the following steps

[0110] S51. Calculate the cosine similarity between the target object feature vector and the enhanced search frame features to obtain the similarity map between the target object and the search image

[0111] S = f obj · F s

[0112] Each element S in the similarity map S (c,i,j) (0 < i < H, 0 < j < W) represents the similarity degree between the target object feature vector and each image patch of the search frame. The larger the value of S (C,i,j) , the higher the similarity, that is, this position can be considered as the foreground point in the video object tracking task; conversely, S (C,i,j)The smaller the value, the lower the similarity, that is, this position can be considered as a background point in the video object tracking task, and S is selected (C,i,j) The point on the image corresponding to the maximum value and the point on the image corresponding to the minimum value are used as a pair of positive and negative position point pairs, denoted as P maxmin =(P max ,P min );

[0113] S52. Using the positive and negative position point pair P maxmin as a cue, decode the enhanced search frame feature to output the target-oriented search frame feature That is:

[0114]

[0115] S53. Using the target-oriented search frame feature obtain the target classification score map local offset map normalized size map of the target through a fully convolutional network. The center point of the target object corresponds to the highest score in the classification map. Based on the target classification score map, local offset map, and normalized size map, the initial target object bounding box bbox 0 is:

[0116]

[0117]

[0118] Therefore, the initial target object bounding box bbox 0 is:

[0119] bbox 0 =(x, y, w, h)

[0120] =(x d +O(0, x d , y d ), y d +O(1, x d , y d ), S(0, x d , y d ), S(1, x d , y d ))

[0121] bbox 0 =xywh2xyxy(bbox 0 )

[0122] Among them, xywh2xyxy(·) means converting the bounding box of [center point x coordinate, center point y coordinate, target box width w, target box height h] into the format of [upper left corner point x coordinate, upper left corner point y coordinate, lower right corner point x coordinate, lower right corner point y coordinate].

[0123] S54. Using the preliminary target object bounding box bbox 0 As a hint, input the target-guided search frame feature into the mask decoder of the hint-enabled segmentation model. At this time, when decoding, instead of outputting the segmentation result, the final search frame feature is output That is:

[0124]

[0125] S55. Similar to step S53, using the final search frame feature obtain the target classification score map of the target through a fully convolutional network local offset map normalized size map Obtain the final target object bounding box bbox based on the target classification score map, local offset map, and normalized size map, achieving the effect of self-optimization:

[0126]

[0127]

[0128] bbox = (x, y, w, h)

[0129] = (x d '+ O'(0, x d ', y d '), y d '+ O'(1, x d ', y d '), S'(0, x d ', y d '), S'(1, x d ', y d '))

[0130] bbox = xywh2xyxy(bbox).

[0131] Step S6: Model training. Under the video single-object tracking encoder-feature enhancement-decoder paradigm, construct a single-object tracking model that includes an encoder based on a promptable segmentation model, a self-attention enhancement unit for template and search region features, a target-oriented foreground prompt unit, and a self-optimizable target box decoder, and train the single-object tracking model on the server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges.

[0132] Specifically, it mainly includes the following steps:

[0133] S61. Use the server to perform image cropping and data augmentation. Divide the single-object tracking dataset (LaSOT dataset, TNL2K dataset, OTB-Lang dataset) that has been annotated into a training set and a test set according to the official method. After reading each frame of the picture, perform cropping and data augmentation operations on it. Among them, the cropping method is: crop a rectangular image centered on the area where the target is located, and the length and width of this rectangular image are 2 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the original video boundary is filled with the pixel average value, and finally scale this rectangular image to 256×256 or 384×384 to form a template frame image; crop a rectangular image centered on the area where the target is located, and the length and width of this rectangular image are 4 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the original video boundary is filled with the pixel average value, and finally scale this rectangular image to 256×256 or 384×384 to form a search frame image. When performing data augmentation, flip the image horizontally with a probability of p = 0.5, grayscale the image with a probability of p = 0.05, and perform center point and size jitter on the image with a probability of p = 0.2;

[0134] S62. Use the server to execute step S2, that is, based on the encoder of the promptable segmentation model, extract features from both the search frame image and the template frame image; then execute step S3, and perform feature interaction and enhancement on the template frame image features and the search frame image features through the self-attention enhancement unit for template and search region features; then execute step S4 to obtain target-oriented foreground prompts; finally execute step S5, and obtain the target classification score map, local offset map, and normalized size map of the tracked target through a self-optimizable target box decoder, and then the target object bounding box can be obtained;

[0135] S63. Use the server to train the network model. The present invention uses a promptable segmentation model as the basic model. In addition, the present invention uses two ways of pairing image resolutions, namely template frame (256×256), search frame (256×256) and template frame (384×384), search frame (384×384).

[0136] During training, an end-to-end training method is adopted, and the overall expression of the loss function is as follows:

[0137]

[0138] Among them, L cls represents the focal loss for classification; L iou and L 1 are both used to supervise bounding box regression, representing the generalized intersection over union loss and the L1 loss respectively. λ iou and represent the weights of the generalized intersection over union loss and the L1 loss respectively, and are set to: λ iou = 2,

[0139] S64. Train the single-object tracking model on the server, set the learning rate to 4×10 -4 , the batch size to 16, and train for a total of 100 epochs. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges to obtain the locally optimal network parameters.

[0140] Step S7. Model inference. Use the trained network model to track a specified single object in the video sequence to be tracked.

[0141] Specifically, it mainly includes the following steps:

[0142] S71. For a given video, initialize the first frame of the video as the template frame for video object tracking, and given the bounding box of the template frame. According to this bounding box, divide a region from the template frame with a preset size change as the template frame image, and the video object tracking network starts tracking from the second frame;

[0143] S72. For each frame from the second frame and later, using the center point of the target box in the previous frame as a reference, divide a region with a preset size and distance as the search frame image, and input the template frame image and the search frame image into the network model to output the tracking result of the current frame.

[0144] A video object tracking method based on a promptable segmentation model provided by the present invention applies the promptable segmentation model well to the video object tracking task, taking into account the modeling of the overall information of the template image and the specific information of the tracking object. The tracking object features provide important prompt information for the target box decoding stage of the search frame, greatly solving the adverse impact of the background information in the template image on this task and improving the robustness of the tracking. In addition, the target box decoder has the ability of self-optimization, which can timely feedback and optimize the output bounding box, improving the accuracy of the tracking. The present invention can finally accurately and stably track the target in many difficult actual scenarios, achieving better object tracking effects compared with other methods.

[0145] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.

Claims

1. A video object tracking method based on a promptable segmentation model, characterized in that, it includes the following steps: S1. Construct a video single-object tracking encoder-feature enhancement-decoder paradigm; S2. Construct an encoder based on a promptable segmentation model; S3. Construct a self-attention enhancement unit for template and search region features; S4. Construct a target-oriented foreground prompt unit; S5. Construct a self-optimizable object box decoder; S6. Under the video single-object tracking encoder-feature enhancement-decoder paradigm, construct a single-object tracking model including an encoder based on a promptable segmentation model, a self-attention enhancement unit for template and search region features, a target-oriented foreground prompt unit, and a self-optimizable object box decoder, and train the single-object tracking model on a server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges; S7. Use the trained network model to track a specified single object in the video sequence to be tracked; The specific steps of step S1 include the following steps: S11. An image encoder based on a promptable segmentation model, combined with an adapter, a commonly used tool in the field of natural language processing, to establish an image feature extractor, and input the search frame and the template frame respectively extract features with the image feature extractor to obtain the search frame feature x s and the template frame feature x t ; S12. Take the search frame feature x s and the template frame feature x t and use the self-attention enhancement unit of the template and the search area features for efficient fusion to obtain the enhanced search frame feature F s and the enhanced template frame feature F t ; S13. In the target-oriented foreground prompting unit, the enhanced template frame feature F is used t , and segmentation is performed through a promptable segmentation model to obtain a segmentation mask of the object to be tracked in the video object tracking task. Using the segmentation mask, the feature f of the target object is obtained obj ; S14. Input the feature f of the target object obj and the enhanced search frame feature F s into a self-optimizable target box decoder to obtain a tracking result, where the tracking result includes the target classification score map P, the local offset map O, and the normalized size map S of the tracked target.

2. According to the video object tracking method based on a promptable segmentation model described in claim 1, characterized in that, the specific steps of step S2 include the following steps: S21. Retrain the absolute position encoding and relative position encoding of the vision Transformer in the image encoder of the promptable segmentation model to match the image resolutions of 256×256 and 384×384; S22. Use an adapter to cooperate with the image encoder in the promptable segmentation model as an image feature extractor, and extract features from the input search frame and the template frame respectively with this image feature extractor to obtain the search frame feature and the template frame feature That is: Among them, the role of PatchEmbed(·) is to map image patches to the hidden space, is the absolute position encoding, LN(·) is the layer normalization, MSA(·) is the multi-head attention mechanism, Adapter(·) is the adapter, and num_block represents the number of unit blocks in the image feature extractor.

3. According to the video object tracking method based on a promptable segmentation model described in claim 2, characterized in that, The adapter consists of a fully connected downsampling layer, an activation layer, a fully connected upsampling layer, and a residual connection.

4. According to the video object tracking method based on a promptable segmentation model described in claim 1, characterized in that, The self-attention enhancement unit for template and search region features consists of 2 self-attention units, and each self-attention unit consists of a multi-head attention layer (MSA), layer normalization (LN), a multi-layer perceptron (MLP), and a residual connection.

5. According to the video object tracking method based on a promptable segmentation model described in claim 4, characterized in that, the specific steps of step S3 include the following steps: S31. Concatenate the search frame features obtained in step S22 and the template frame features to form and add absolute position encoding Then send it to two self-attention units for processing; S32. Send the splicing feature to the self-attention enhancement unit of the template and the search area feature, and segment the obtained result to obtain the enhanced search frame feature and the enhanced template frame feature 6. According to the video object tracking method based on a promptable segmentation model described in claim 5, characterized in that, The specific implementation process of step S4 is as follows: Using the target box of the template frame as a prompt, extract features using the prompt encoder in the promptable segmentation model According to the features prompted by the target box and the enhanced template frame features Use the mask decoder in the promptable segmentation model to obtain the segmentation mask m of the target object t , then utilize the screening ability of the mask to act on the enhanced template frame features, and obtain the target object feature vector through average pooling 7. According to the video object tracking method based on a promptable segmentation model described in claim 6, characterized in that, the specific steps of step S5 include the following steps: S51. Calculate the cosine similarity between the feature vector of the target object and the enhanced search frame features to obtain the similarity map between the target object and the search image S = f obj ·F s Each element S in the similarity graph S (C,i,j) (0 < i < H, 0 < j < W) represents the degree of similarity between the target object feature vector and each image patch in the search frame. S (C,i,j) The larger the value, the higher the similarity, that is, this position can be considered as the foreground point in the video object tracking task; conversely, S (C,i,j) The smaller the value, the lower the similarity, that is, this position can be considered as the background point in the video object tracking task. Select S (C,i,j) The points on the image corresponding to the maximum value and the points on the image corresponding to the minimum value as a pair of positive and negative position point pairs, denoted as P maxmin =(P max , P min ); S52. Use the positive and negative position point pair P maxmin as a hint to decode the enhanced search frame feature and output the target-oriented search frame feature That is: S53. Using the target-oriented search frame features Obtaining a target classification score map of the target through a fully convolutional network Local offset map Normalized size map Among them, the center point of the target object corresponds to the one with the highest score in the classification map; obtaining a preliminary target object bounding box bbox based on the target classification score map, the local offset map, and the normalized size map 0 ; S54. Use the preliminary target object bounding box bbox 0 As a hint, input the target-guided search frame feature into the mask decoder of the hintable segmentation model, and output the final search frame feature S55. Utilize the final search frame features Obtain the target classification score map of the target through a fully convolutional network Local offset map Normalized size map Obtain the final target object bounding box bbox based on the target classification score map, the local offset map, and the normalized size map, achieving the effect of self-optimization.

8. According to the video object tracking method based on a promptable segmentation model described in claim 7, characterized in that, the specific steps of step S6 include the following steps: S61. Image cropping and data augmentation using a server: The cropping method is as follows: A rectangular image is cropped with the area where the target is located as the center. The length and width of this rectangular image are 2 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the boundary of the original video is filled with the pixel average value. Finally, this rectangular image is scaled to 256×256 or 384×384 to form a template frame image. A rectangular image is cropped with the area where the target is located as the center. The length and width of this rectangular image are 4 times the length and width of the target rectangular box. The part of the rectangular box that exceeds the boundary of the original video is filled with the pixel average value. Finally, this rectangular image is scaled to 256×256 or 384×384 to form a search frame image. Data augmentation is performed on the template frame and the search frame, including random inversion, random grayscaling, center point position jittering, and bounding box size jittering operations. S62. Use the server to execute step S2, that is, based on the encoder of the promptable segmentation model, perform feature extraction on both the search frame image and the template frame image. Then execute step S3, and perform feature interaction and enhancement on the template frame image feature and the search frame image feature through the self-attention enhancement unit of the template and search area features. Then execute step S4 to obtain the target-oriented foreground prompt. Finally, execute step S5, and obtain the target classification score map, local offset map, and normalized size map of the tracked target through a self-optimizable target box decoder, and then obtain the target object bounding box. S63. Use the server to train the network model, and perform training in an end-to-end manner. The overall expression of the loss function is: Among them, L cls represents the focal loss for classification; L iou and L 1 are both used to supervise the bounding box regression, representing the generalized intersection over union loss and the L1 loss respectively, λ iou and represent the weights of the generalized intersection over union loss and the L1 loss respectively, and are set to: λ iou = 2, S64. Train the single-object tracking model on the server. By reducing the overall loss value of the network loss function, optimize the network parameters until the network converges to obtain the locally optimal network parameters.

9. A video object tracking method based on a promptable segmentation model according to claim 8, wherein, the specific steps of step S7 are as follows: S71. For a given video, initialize the first frame of the video as the template frame for video object tracking, and given the bounding box of the template frame. According to this bounding box, divide a region with a preset size change from the template frame as the template frame image. The video object tracking network starts tracking from the second frame. S72. For each frame starting from the second frame, use the center point of the target box in the previous frame as a reference, and divide a region with a preset size and distance as the search frame image. Input the template frame image and the search frame image into the network model to output the tracking result of the current frame.

Citation Information

Patent Citations

  • Transform-based online update target tracking method and system

    CN114998601A

  • Unmanned aerial vehicle target tracking method based on mask pre-training

    CN115393396A