A Dynamic Inference Path Object Tracking Method Based on Conditional Early Exit Mechanism
By introducing a conditioned early retreat mechanism into the target tracking algorithm, dynamically adjusting the inference path, the problem of limitations in computing power in different application scenarios and equipment is solved, and efficient and flexible target tracking is achieved.
Patent Information
- Application Number
- CN202211532182.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-12-01
AI Technical Summary
The existing target tracking algorithms are limited by different application scenarios and device computing capabilities during deployment, making it difficult to achieve efficient target tracking while ensuring performance.
The conditioned early retreat mechanism is adopted to dynamically adjust the inference path, and adaptively select the inference path according to the difficulty of inputting the video sequence sample to achieve a trade-off between performance and speed.
While ensuring high accuracy of tracking results, dynamically selecting different inference paths greatly saves calculations, improves the actual speed of tracking methods, and can be flexibly deployed on different computing power devices.
Smart Images

Figure CN115861374B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning and visual object tracking, and relates to a visual transformer feature extraction model, a conditional early exit mechanism, and a dynamic inference network. Background Art
[0002] Visual object tracking is an important branch in the field of computer vision (CV). Its main task is as follows: Given an arbitrary target in the initial frame, the algorithm will accurately track the position of the target in subsequent video frames, predict the scale of the target, and output the bounding box of the target, so as to realize the understanding of the target's existence, position, motion state and other behaviors, and support the completion of more advanced tasks. The Siamese series of tracking algorithms have attracted much attention. Among them, the SiamFC tracker (Fully-Convolutional Siamese Networks for Object Tracking) proposed by Luca et al. is a representative of the pioneer work of this series of methods. In recent years, especially with the introduction of more powerful feature extraction backbone networks, the application of attention mechanisms, the injection of online training modules and cascade structures, the progress of object tracking algorithms has been further promoted. The video sources for object tracking can be video clips obtained from surveillance cameras, live video streams of sports competitions, etc. Object tracking algorithms mainly include steps such as video frame extraction and preprocessing, search area extraction, object feature extraction and interactive tracking, and regression prediction of object positions. Among them, object feature extraction and interaction, and object position regression prediction are the core contents, mainly for continuously tracking objects of any category, any motion state, and in any scenario. The performance of an object tracking algorithm is mainly measured by accuracy and speed.
[0003] The early exit mechanism of neural networks aims to dynamically allocate computing components and computing paths during network inference. Such methods usually adaptively allocate different network layers, sub-networks, and inference paths for each different input sample, which is a major research direction in the field of efficient neural network computing. The early exit mechanism has recently been applied to image recognition tasks, and its main implementation method is to introduce cascaded intermediate state classifiers in the backbone network. For example, in order to save computational effort, Xin et al. developed SkipNet (Skipnet: Learning dynamic routing in convolutional networks), which explores selectively skipping some residual modules for image classification according to different input samples based on ResNet. Andreas et al. proposed ConvNet-AIG (Convolutional networks with adaptive inference graphs), in which a gating function is used to decide whether to execute or skip at each residual layer, thereby realizing the adjustment of the computing architecture based on the input image. Generally, the judgment condition for the early exit action is the difficulty of the input sample, and further dynamic path inference analysis is carried out based on this. However, current research work on the early exit mechanism mainly focuses on basic image classification tasks, and there is still a lack of exploration on how to effectively introduce this idea into more complex downstream tasks required by practical applications. The present invention applies the conditional early exit mechanism to the field of video object tracking with spatio-temporal dimensions, effectively improving the efficiency of the object tracking algorithm during actual operation. Summary of the Invention
[0004] The present invention aims to provide a high-efficiency and high-performance dynamic inference path object tracking method based on the conditional early exit mechanism, to solve the problem that existing object tracking algorithms are restricted by the computing capabilities of different application scenarios and different operating devices during deployment. The present invention can adaptively adjust the inference path according to the difficulty of the input video sequence samples to both meet the performance requirements and achieve efficient object tracking. At the same time, the method proposed by the present invention only needs to train one model on the training data, and for devices with different computing performances, the trade-off between performance and speed can be achieved by only adjusting one decision parameter. The method described in the present invention can be deployed on various practical application platforms such as monitoring devices or the on-board control terminal of drones to act as a visual module for real-time tracking and analysis of objects.
[0005] The technical solution of the present invention is as follows:
[0006] A dynamic inference path object tracking method based on the conditional early exit mechanism, the steps are as follows:
[0007] Step 1: Obtain continuous video frames to be processed by means of an imaging device such as a camera;
[0008] Step 2: Input a continuous video stream, and at the same time specify any target to be tracked in the initial frame of the video;
[0009] Use the vector B0 to represent the position and size of the initial target:
[0010]
[0011] where, is the position of the center point of the initial target, and (h 0 , w 0 ) is the scale of the initial target.
[0012] Step 3: Generate a template region according to the specified target to be tracked. The template region is an outward expansion region of the target bounding box, with its center position unchanged, and the scale is the geometric mean of γ tem times the target scale (h 0 , w 0 ). Here, a γ tem value of 2 is adopted. At the same time, based on the given initial target, generate a search region for the frame to be tracked. According to the continuity of the target motion trajectory, the center position of the search region is the same as the center position of the target in the previous frame (if the previous frame is the initial frame, the center position is the center position of the target specified in the initial frame), and the scale of the search region is the geometric mean of γ sea times the scale of the target in the previous frame. Here, a γ sea value of 4 is adopted.
[0013] Step 4: Extract the depth features of the template and the search region through the encoder layer of the transformer;
[0014] The encoder layer of the Transformer is taken from the ViT model. A single transfomer encoding layer is mainly composed of structures such as a multi-head attention module (MSA), layer normalization (LN), a feed-forward network (FFN), and a residual connection; the multi-head attention module receives a token input with a dimension of 768. This module first calculates three new matrices: Query, Key, and Value. These three matrices are obtained by multiplying the input token by a randomly initialized matrix. Query and Key are multiplied, multiplied by a scaling constant, then a softmax operation is performed, and finally multiplied by the Value matrix to obtain the self-attention result. The multi-head attention mechanism splits the above process of calculating self-attention into 12 times, and then concatenates all the results as the output of the multi-head attention module; the feed-forward network is mainly composed of a fully connected layer, a GELU activation function, a Dropout layer, a fully connected layer, and a Dropout layer in sequence;
[0015] The steps for the Transformer encoder layer to extract template and search region features are as follows:
[0016] (4.1) Input end processing: Crop, pad, scale, and perform other transformations on the template and search region image patches to make the image size consistent with the network input size;
[0017] (4.2) The image patches pass through the Embedding layer to generate a token sequence. The Embedding layer uses a convolutional layer with 768 convolutional kernels, with a size of 16×16 and a stride of 16. Then, corresponding position encodings are added to the generated template and search region and and the template and search region tokens are concatenated:
[0018]
[0019] (4.3) The concatenated template and search region features H pass through N stacked transformer encoder layers to generate deep features H N ;
[0020] Step 5: The encoded deep features H N During the backpropagation process, they pass through path decision nodes, and each decision node E i will judge the current target discrimination state;
[0021] The dynamic path inference process specifically includes the following steps:
[0022] (5.1) Use the stacked transformer encoder layers in step (4.3) as the backbone network. When the encoded features extracted from the backbone network encounter decision points, they enter the adaptation layer, which is composed of transformer encoding layers, and its initial parameters are loaded from the corresponding network layers in the backbone network in step (4.3). Specifically, the first set of adaptation layer parameters is loaded from the parameters of the 3rd - 4th layers of the backbone network, and the second set of adaptation layer parameters is loaded from the parameters of the 7th layer of the backbone network; the adaptation layer for the first decision point has 2 layers, the adaptation layer for the second decision node has 1 layer, and the number of adaptation layers for the third decision point is 0 layer;
[0023] (5.2) The encoded features go through layer normalization and are fed into a bottleneck module to map their dimensions from 768 dimensions to 256 dimensions;
[0024] (5.3) At this time, the encoded features are fed into the IoU score prediction head to predict the IoU score of the current node. The predicted IoU score will be used as the judgment condition for early exit. The IoU score prediction head consists of a 3-layer MLP. The first layer maps the 256-dimensional input to 512 dimensions, the middle layer maintains 512 dimensions unchanged, and the third layer maps the 512-dimensional feature sequence to a 1-dimensional IoU score;
[0025] Step 6: Decision condition judgment. Determine whether the early exit condition of the model is met by the level of the IoU score value. According to the computing power of the actual deployment platform and the demand for algorithm speed in the instance application scenario, different IoU thresholds τ can be set. If the early exit condition is met, the encoded features will exit after passing through the prediction head from the current decision node, that is, the target tracking process of the current frame is completed; in the case where it is determined that the early exit condition has not been met, the deep encoded features will continue to propagate backward until they reach the last layer of the backbone network. During this period, the encoded features of the previous decision points will be reused by the subsequent nodes;
[0026] The decision-making process of the conditional early exit mechanism is as follows:
[0027] (6.1) Compare the IoU score score obtained in step (5.3) with the dynamic network set threshold τ. If score ≥ τ, the early exit condition is met, and the encoded features will directly enter the corner prediction head to predict the upper left and lower right corner points of the target location, and finally output the target location and scale of the current frame:
[0028]
[0029] The corner prediction head consists of 4 RepVGG blocks and a 3×3 convolutional layer. The feature dimension is sequentially mapped from 256 dimensions to 128, 64, 32, 16, and 2 layers. The last two layers of feature maps represent the prediction maps of the upper left and lower right corner points respectively. The highest response points of the corner prediction maps are used as the upper left and lower right corner points of the target prediction, and the final predicted target bounding box is generated;
[0030] (6.2) Compare the IoU score score obtained in step (5.3) with the dynamic network set threshold τ. If score < τ, the early exit condition is not met. The encoded features at the decision point will continue to propagate backward from the backbone network until the next decision point is encountered; the encoded features generated in step (5.2) will be reused in the subsequent decision network, and the reuse method is to directly add them to the feature encoding here; then go through the links in steps (5.2) and (5.3);
[0031] Step 7: Sequentially pass through each early exit decision point. If the early exit condition is met, predict the location and scale of the target in the current frame and end the prediction for this frame. If the condition is not met, continue to propagate backward through the subsequent decision nodes to obtain the final prediction result for the current frame. Predict the input video frames sequentially to obtain the target tracking results for all frames of the corresponding video sequence.
[0032] Advantages of the present invention:
[0033] (1) The feature extraction backbone of the target tracker uses a ViT model pre-trained with MAE. At the same time, multiple early exit decision points are set at different encoder layers for dynamic path inference. While ensuring the high accuracy of the tracking results, different inference paths are dynamically selected for different video frames, greatly saving the computational effort for inference on simple sample frames and improving the actual speed of the tracking method.
[0034] (2) The present invention can be deployed on devices with different computing powers. Only one training is required, and the decision boundary can be flexibly set according to conditions such as the computing power of the edge device in the subsequent actual deployment to meet the trade-off between algorithm performance and speed in practical applications. Description of the Drawings
[0035] Figure 1 It is a schematic diagram of the algorithm flow of the backbone tracker.
[0036] Figure 2 It is a flowchart of the conditional early exit dynamic path tracking algorithm. Detailed Embodiment
[0037] The following further illustrates the detailed embodiment of the present invention in combination with the drawings and technical solutions.
[0038] Figure 1It is the algorithm flowchart of the backbone tracker, and its main body consists of three major parts: feature embedding, feature extraction, and target prediction. In the feature embedding part, the preprocessed template and the sliced images in the search area pass through a convolutional sliding window with a downsampling factor of 16 and a stride of 16, and are mapped to a feature embedding with a dimension of 768. At the same time, position encoding is added to the embeddings of the template and the search area and concatenated with the IoU encoding to form a feature encoding sequence for backpropagation; then, three groups of a total of 12-layer transformer encoding layers are used to extract deep features, where each layer of the transformer consists of a multi-head attention module (MSA), layer normalization (LN), a feed-forward network (FFN), residual connections, and other structures; the target prediction part is mainly composed of a bounding box prediction head and an IoU prediction head. After the deep features are extracted, the feature encoding of the search area and the IoU feature encoding are respectively input into the bounding box prediction network and the IoU prediction network. The bounding box prediction network generates two corner prediction maps, and the upper left and lower right corner points of the target are obtained according to the highest response. The IoU prediction network inputs the predicted IoU score of the target in the current frame.
[0039] Figure 2 It is the algorithm flowchart of the conditional early-exit dynamic path tracking algorithm. In Figure 1 the structure of the complete network, early-exit nodes and a dynamic inference network are added. The parameters of the backbone part of the dynamic path tracking algorithm are directly loaded from the basic tracker. When processing the video stream, in the early stage, Figure 1 it is the same as the process. First, the template and the sliced search area are subjected to feature embedding, and then sent to the subsequent feature extraction network. An early-exit decision point is added to the feature extraction network here. Therefore, including the end output port, there are a total of 3 exits for feature output, which are respectively set at the 2nd, 6th, and 12th layers of feature extraction. When encountering the first early-exit decision point, the encoded features are sent to the first early-exit decision network. The decision network predicts the predicted IoU of the target at this node and compares it with the set early-exit prediction. If the early-exit condition is met, it directly enters the target bounding box prediction module, outputs the final target tracking result, and ends the prediction of this frame; if the early-exit condition is not met, the encoded features start from the current node of the backbone network, continue to propagate backward, and enter the next decision network again. Before entering, the encoded features extracted from the previous early-exit decision network are reused, and the subsequent links are the same as those described above until the exit condition of a certain early-exit node is met or it propagates to the last layer, ending the target prediction of the current frame.
[0040] The training data of the tracking algorithm model of the present invention consists of the training sets of GOT-10k, COCO, LaSOT, and TrackingNet. When generating training sample pairs, data augmentation methods such as random position jitter and random scale jitter are used for data augmentation. The optimizer selects Adamw, and the weight decay coefficient is set to 10 -4, the initial learning rate of the feature extraction part is set to 4×10 -5 , and the initial learning rate of other network structures is set to 4×10 -4 . The backbone network and the dynamic path model are trained for 300 epochs respectively. The learning rate decays by 10 times at 240 epochs. The dynamic path model loads the parameters of the trained backbone network as initialization parameters. In the inference process, the input image size of the network is 128×128 for the template and 256×256 for the search area. The corner prediction head directly obtains the bounding box results of the target without additional post-processing operations.
Claims
1. A dynamic inference path object tracking method based on a conditional early exit mechanism, characterized in that, The steps are as follows: Step 1: Obtain a continuous video stream to be processed with the help of an imaging device; Step 2: Input the continuous video stream and specify the initial target to be tracked in the initial frame of the video at the same time; Use the vector B0 to represent the position and size of the initial target: Among them, is the position where the initial target center point is located, (h 0 , w 0 ) is the scale of the initial target; Step 3: Generate a template region based on the specified initial target to be tracked. The template region is an outward expansion region of the initial target bounding box, with the same center position and a scale of γ tem times the geometric mean of the scale of the initial target (h 0 , w 0 ); meanwhile, generate a search region for the frame to be tracked based on the given initial target. According to the continuity of the target motion trajectory, the center position of the search region is the same as the target center position of the previous frame; if the previous frame is the initial frame, the center position is the target center position specified in the initial frame; the scale of the search region is the geometric mean of γ sea times the scale of the target in the previous frame; Step 4: Extract the depth features of the template region and the search region through the encoder layer of the transformer; The encoder layer of the Transformer is taken from the ViT model. A single Transformer encoding layer is mainly composed of a multi-head attention module, layer normalization, a feed-forward network, and a residual connection; the multi-head attention module receives token inputs with a dimension of 768, and first calculates three new matrices: Query, Key, and Value; the three new matrices are obtained by multiplying the input tokens by a randomly initialized matrix; the Query matrix and the Key matrix are multiplied, multiplied by a scaling constant, then a softmax operation is performed, and finally multiplied by the Value matrix to obtain the self-attention result; the multi-head attention mechanism splits the process of obtaining self-attention into 12 times, and then concatenates all the self-attention results as the output of the multi-head attention module; the feed-forward network is mainly composed of a fully connected layer, a GELU activation function, a Dropout layer, a fully connected layer, and a Dropout layer connected in sequence; The steps for the encoder layer of the Transformer to extract the features of the template region and the search region include the following: (4.1) Input end processing: Transform the image patches of the template region and the search region to make the image size consistent with the network input size; (4.2) The image patch passes through the Embedding layer to generate a token sequence; the Embedding layer uses a convolutional layer with 768 convolutional kernels, with a size of 16×16 and a stride of 16; then, corresponding positional encodings are added to the generated template region and search region and and the tokens in the template region and search region are concatenated: (4.3) The spliced template region and search region feature H0 generate the depth feature H after passing through the N stacked transformer encoder layers N ; Step 5: Encoded depth feature H N During the backpropagation process, it will pass through the path decision nodes, and each decision node E i will judge the current target discrimination state; The dynamic path inference process specifically includes the following steps: (5.1) Use the stacked Transformer encoder layers in step (4.3) as the backbone network. The encoded features extracted in the backbone network enter the adaptation layer when encountering decision points. The adaptation layer is composed of Transformer encoding layers, and its initial parameters are loaded from the corresponding network layers in the backbone network in step (4.3). Specifically, the first group of adaptation layer parameters is loaded from the parameters of the 3rd - 4th layers of the backbone network, and the second group of adaptation layer parameters is loaded from the parameters of the 7th layer of the backbone network; the adaptation layer of the first decision point has 2 layers, the adaptation layer of the second decision node has 1 layer, and the adaptation layer of the third decision point has 0 layers; (5.2) The encoded features go through layer normalization and are sent into a bottleneck module to map their dimension from 768 dimensions to 256 dimensions; (5.3) At this time, the encoded features are sent into the IoU prediction module, that is, the IoU score prediction head, to predict the IoU score of the current node. The predicted IoU score will be used as the judgment condition for whether to choose early exit; the IoU score prediction head is composed of a 3-layer MLP. The first layer maps the 256-dimensional input to 512 dimensions, the middle layer keeps 512 dimensions unchanged, and the third layer maps the 512-dimensional feature sequence to a 1-dimensional IoU score; Step 6: Decision condition judgment; Determine whether the early exit condition of the model is met based on the IoU score value; Set different IoU thresholds τ according to the computing power of the actual deployment platform and the demand for algorithm speed in the instance application scenario; If the early exit condition is met, the encoded features will exit after passing through the IoU score prediction head from the current decision node, that is, the target tracking process of the current frame is completed; In the case where it is determined that the early exit condition has not been met, the encoded features will continue to propagate backward until they reach the last layer of the backbone network, and the encoded features of the previous decision points will be reused in the subsequent nodes; The decision-making process of the conditional early exit mechanism is as follows: (6.1) Compare the IoU score score obtained in step (5.3) with the dynamic network set IoU threshold τ. If score ≥ τ, the early exit condition is met, and the encoded features will directly enter the corner prediction head to predict the upper left and lower right corner points of the target location, and finally output the target position and scale of the current frame: The corner prediction head consists of 4 RepVGG blocks and a 3×3 convolutional layer. The feature dimension is sequentially mapped from 256 dimensions to 128, 64, 32, 16, and 2 layers. The last two feature maps represent the prediction maps of the upper left and lower right corner points respectively. The highest response points of the corner prediction maps are used as the upper left and lower right corner points of the target prediction, and the final predicted target bounding box is generated; (6.2) Compare the IoU score score obtained in step (5.3) with the dynamic network set threshold τ. If score < τ, the early exit condition is not met; The encoded features at the decision point will continue to propagate backward from the backbone network until the next decision point is encountered; The encoded features generated in step (5.2) will be reused in the subsequent decision network, and the reuse method is to directly add them to the feature encoding here; Then go through the links in steps (5.2) and (5.3); Step 7: Pass through each early exit decision point in turn. If the early exit condition is met, predict the position and scale of the target in the current frame and end the prediction of this frame. If the condition is not met, continue to propagate backward, pass through the subsequent decision nodes, and obtain the final prediction result of the current frame; Predict the input video frames in turn to obtain the target tracking results of all frames of the corresponding video sequence.