Single-object Tracking Method Based on Deep Learning Hierarchical Video Attention
By introducing a layered video attention mechanism and global inter-sliding window causal attention, the robustness of the single-object tracking method in complex scenarios of occlusion and appearance changes is solved, and efficient and real-time single-object tracking is achieved, which improves the adaptability and training efficiency of the model.
Patent Information
- Application Number
- CN202510412123.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing single-objective tracking method is not robust enough in scenarios of complex occlusion and appearance changes, making it difficult to achieve real-time high-precision tracking under limited computing resources, and traditional methods are inefficient in training and inference scenarios.
The hierarchical video attention mechanism is adopted, and by constructing a video sequence as the model input sample, a coordinate space perturbation strategy and a search area cropping module are introduced, combining hierarchical feature extraction and causal attention in the global inter-frame sliding window to achieve consistency optimization of training and reasoning, and the model adaptability is improved through data augmentation and autoregression tracking.
It realizes high-precision real-time tracking in complex video scenarios, improves the anti-interference performance and training efficiency of the model, breaks through the limitations of traditional methods in timing dependence and frame-level context modeling, and has long-range context reasoning capabilities.
Smart Images

Figure CN119919457B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a single-object tracking method based on hierarchical video attention of deep learning, and belongs to the technical field of computer vision and image processing. Background Art
[0002] With the continuous development of artificial intelligence and computer vision technologies, single-object tracking has wide application values in fields such as security monitoring, driverless, robot navigation, and human-computer interaction. Traditional single-object tracking methods are mostly based on correlation filtering or template matching. By giving the target position in the initial frame and extracting its features, similar feature regions are continuously searched in subsequent frames to estimate the target position. However, these methods face several challenges in practical applications.
[0003] Occlusion and interference: When the target is partially or completely occluded, traditional methods based on template matching or correlation filtering usually rely on fixed update mechanisms or manually designed occlusion handling strategies, and it is difficult to achieve accurate and robust tracking in complex scenarios.
[0004] Appearance changes: The target will undergo scale changes, pose changes, and illumination changes, etc. in the real scene. Traditional methods often require manually designing features or using models based on limited training data, and it is difficult to cope with the appearance diversity of the target.
[0005] Balance between real-time performance and accuracy: The single-object tracking task often requires real-time processing with limited computing resources, and has high requirements for the efficiency of the algorithm. However, how to ensure the accuracy of tracking while improving the tracking speed has always been a difficult problem to balance.
[0006] In recent years, deep learning has demonstrated powerful feature learning capabilities and robustness in computer vision tasks such as object detection, semantic segmentation, and video understanding. Introducing it into single-object tracking can effectively improve the adaptability to complex scenarios. Among them, tracking algorithms based on deep convolutional networks or Siamese structures have made remarkable progress. However, in dealing with long-sequence videos, strongly occluded scenarios, and drastic changes in the target appearance, the robustness and stability of the existing technologies still need to be further improved. At the same time, with the successful application of new network structures such as Transformer in the field of images and videos, researching how to combine the hierarchical attention mechanism with single-object tracking and using local and global attention to balance speed and accuracy has become a hot and difficult point in the current technological development. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a single-object tracking method based on deep learning hierarchical video attention. The goal is to output the coordinates of the object in all subsequent frames by inputting a video and the coordinates of the object in the first frame. By flexibly applying the hierarchical attention mechanism in the network structure, while maintaining a high tracking accuracy, taking into account real-time performance and adaptability to complex video scenarios, so as to meet diverse application requirements.
[0008] To achieve the above object, the technical solution of the present invention is as follows. A single-object tracking method based on deep learning hierarchical video attention, the method comprising the following steps:
[0009] Step 1, read the data set to obtain the original video data, the object coordinates and the object occlusion ratio in each frame.
[0010] Step 2, use the object occlusion ratio to screen out the reference frame, and construct an initial video sequence with the reference frame as the first frame and an initial coordinate ground truth sequence.
[0011] Step 3, perform coordinate perturbation on the initial coordinate ground truth sequence, use the perturbed coordinates to crop the initial video sequence to obtain a cropped video sequence, and use the perturbed coordinates to transform the initial coordinate ground truth sequence to obtain a cropped coordinate ground truth sequence.
[0012] Step 4, perform data augmentation on the cropped video sequence and the cropped coordinate ground truth sequence to obtain an augmented video sequence and an augmented coordinate ground truth sequence.
[0013] Step 5, use the augmented video sequence and the augmented coordinate ground truth sequence to train the hierarchical inter-frame sliding window causal attention network model.
[0014] Step 6, use the trained hierarchical inter-frame sliding window causal attention network model for autoregressive object tracking.
[0015] Among them, in step 1, the video sampling frequency of the data set is set to 10 frames per second.
[0016] Among them, in step 2, when the occlusion ratio is greater than 50% and there are at least frames from the end frame of the video, it is determined as a valid frame, where is a natural number greater than or equal to 48; when randomly selecting a valid frame from all valid frames as the reference frame, sampling optimization is performed based on the variance of the occlusion ratio within the interval.
[0017] Specifically, the window length is set to frames, traverse the occlusion ratio of all frames, calculate the variance within its corresponding window, screen out the variance calculation weights at the positions of the valid frames, and normalize the weights to form a probability distribution:
[0018] ,
[0019] Among them, represents the probability that the th valid starting frame is selected, represents the variance of the occlusion ratio of the window corresponding to the th valid starting frame, represents the variance of the occlusion ratio of the window corresponding to the th valid starting frame;
[0020] Calculate the probability distribution and select the final reference frame using the random weighted sampling method;
[0021] The initial video sequence is defined as , among which, represents the th frame starting from the reference frame, and the initial coordinate true value sequence , among which, represents the initial coordinate true value of the th frame.
[0022] Among them, in step 3, coordinate perturbation uses the dynamic noise injection technique, and the initial coordinate true value of the th frame in the initial coordinate true value sequence is expressed as:
[0023] ,
[0024] Among them, is the abscissa of the upper left vertex of the th frame box, is the ordinate of the upper left vertex of the th frame box, is the width of the th frame box, is the height of the th frame. The initial coordinate true value of the perturbed th frame is calculated as:
[0025] ,
[0026] ,
[0027] ,
[0028] ,
[0029] ,
[0030] Among them, is the The abscissa of the perturbation of the upper left vertex of the frame is the ordinate of the perturbation of the upper left vertex of the frame is the width of the frame perturbation is the height of the frame perturbation represents the natural constant represents the sampling value of Gaussian noise with a variance of 1 represents the sampling value of uniform noise with a value range between 0 and 1;
[0031] The search area cropping adopts an adaptive magnification strategy, with the perturbed coordinates as the center, the cropping side length of the frame is calculated as:
[0032]
[0033] where is the cropping factor for controlling the cropping scale;
[0034] Based on the cropping side length and the perturbed coordinates, a cropped video sequence is generated, and its expression is:
[0035] ,
[0036] where each element , is generated through the following,
[0037] ,
[0038] ,
[0039] represents the operation of performing a square cropping on the image with the center point of as the center coordinate and a side length of ;
[0040] The true value sequence of the cropping coordinates is obtained through the following processing. When generating the cropped video sequence , the initial true value sequence of the coordinates is subjected to an equivalent mapping process to finally form the true value sequence of the cropping coordinates.
[0041] This method addresses the issue of inconsistent training and inference scenarios in existing object tracking technologies and proposes an innovative technical improvement scheme. Different from the traditional single-frame input mode, this method constructs a video sequence as the input sample of the model for the first time, and realizes the consistency optimization of training and inference through the following key technical breakthroughs. First, an innovative coordinate space perturbation strategy is introduced, and by applying random displacement deviations during the training phase, the prediction coordinate errors generated in the actual inference process are effectively simulated. Second, the image preprocessing process in the inference phase is accurately reproduced through the search area cropping module, and an end-to-end training and verification closed-loop is constructed. This dual technical scheme successfully bridges the feature distribution differences between the model training environment and the deployment environment, enables the model to have the ability of multi-frame parallel training, breaks through the limitations of the traditional autoregressive training mode in terms of temporal dependence, and improves the training efficiency by more than 300% compared with the frame-by-frame processing method, providing a new model optimization paradigm for the field of video object tracking.
[0042] Among them, in step 4: enhancing the video sequence Obtained by applying the same rotation angle, flipping direction, Gaussian blur, and color channel permutation spatial transformation to all frames of the cropped video sequence ; enhancing the coordinate ground truth sequence Obtained from the cropped coordinate ground truth sequence After being transformed by the equivalent coordinate mapping when the cropped video sequence is transformed into the enhanced video sequence .
[0043] Among them, in step 5, the hierarchical feature extraction network adopts a four-level progressive architecture. The first two levels use the local intra-frame window attention mechanism for fine-grained feature extraction, and the last two levels use the global inter-frame sliding window causal attention to achieve cross-frame information fusion.
[0044] Among them, the specific calculation steps of the global inter-frame sliding window causal attention module are as follows:
[0045] First, perform feature combination. For the feature input in the th frame, it is calculated as
[0046] ,
[0047] where represents the feature of the th frame picture, represents the feature of the th frame coordinates, represents concatenating the vectors along the first dimension, and then, the input total feature sequence is calculated as:
[0048] ,
[0049] Inter-frame sliding window causal attention restriction matrix The calculation process is as follows:
[0050] ,
[0051] ,
[0052] Among them, represents the attention of the th feature of the th frame input to the th feature of the th frame input. Specifically, represents the total number of sub-blocks of the feature, is a natural constant less than or equal to and greater than 0, which controls that the th frame feature can only focus on the first frame features including itself;
[0053] Finally, the input total feature sequence , the inter-frame sliding window causal attention restriction matrix , and the output feature of the attention module are calculated as:
[0054] ,
[0055] Among them, represents the transpose of the matrix, , and are all learnable parameters of the network model, represents the length of each block of features;
[0056] Rotary position encoding is adopted between frames, and fixed position encoding is adopted within frames; the predicted confidence uses the intersection over union of the predicted coordinates and the predicted ground truth as the ground truth to represent the confidence degree of the model for the current predicted coordinates, which is used to correct the tracking state during inference; the loss function design fuses the coordinate regression error and the predicted confidence error, and uses weighted summation to balance the contribution degrees of different loss terms.
[0057] This method addresses the problem of weak frame-level context modeling ability that is prevalent in existing target tracking technologies, and proposes a breakthrough temporal perception optimization scheme. Compared with the Siamese tracking architecture that relies on fixed template matching or the recursive model limited by memory capacity, this method creatively constructs a global temporal attention network, and for the first time realizes the dynamic correlation modeling of historical features in the entire video domain: through a differentiable attention weight mechanism, an explicit feature interaction channel is established between the current frame and any historical frame, breaking through the double limitations of traditional methods in terms of temporal perception range and modeling depth, enabling the model to have a true long-range context reasoning ability.
[0058] Among them, in step 6, during the inference stage, an incremental processing flow is adopted, and the feature states of the most recent frames are retained through a circular buffer mechanism. In the search area dynamic adjustment algorithm, the search area is initialized with the coordinates of the first frame, and during subsequent predictions, the search area is updated by sliding average according to the intersection over union (IoU) of the predicted coordinates of two adjacent frames as the weight. When the confidence level is lower than the threshold, the range of the search area is expanded.
[0059] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the single-object tracking method based on deep learning hierarchical video attention described above.
[0060] A computer-readable storage medium stores computer instructions thereon. When the computer instructions are executed by a processor, they implement the single-object tracking method based on deep learning hierarchical video attention described above.
[0061] Compared with the prior art, the present invention adopts the above technical solutions and has the following beneficial effects:
[0062] 1. The method of the present invention based on hierarchical video attention realizes real-time online tracking. Even if the object has never been seen before, as long as it is marked in the first frame, it can be continuously tracked in subsequent frames.
[0063] 2. Through the target occlusion ratio screening method, the present invention selects more complex, rapidly changing, and challenging video segments for model training, significantly improving the anti-interference performance of the model in the video stream.
[0064] 3. The method of the present invention based on hierarchical video attention can well handle occlusion situations. Compared with the tracking method based on template matching, it avoids a complex template update mechanism, but can, through the attention mechanism, pay attention to the information of frames a long time ago, enabling the model to still achieve robust tracking according to the information of previous frames when the target is partially or completely occluded.
[0065] 4. The method of the present invention based on hierarchical video attention uses sliding window attention and autoregressive tracking, and can parallelly train the entire video data during training, efficiently utilizing the hardware efficiency. Description of the Drawings
[0066] Figure 1 It is the overall flowchart of the present invention. Detailed Implementation Manner
[0067] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the program flowchart and specific examples.
[0068] Embodiment: A single-object tracking method based on deep learning hierarchical video attention, the method comprising the following steps:
[0069] Step 1, read the data set to obtain the original video data, the target coordinates in each frame, and the target occlusion ratio.
[0070] Step 2, use the target occlusion ratio to screen out the reference frame, and construct an initial video sequence with the reference frame as the first frame and an initial coordinate ground truth sequence.
[0071] Step 3, perform coordinate perturbation on the initial coordinate ground truth sequence, use the perturbed coordinates to crop the initial video sequence to obtain a cropped video sequence, and use the perturbed coordinates to transform the initial coordinate ground truth sequence to obtain a cropped coordinate ground truth sequence.
[0072] Step 4, perform data augmentation on the cropped video sequence and the cropped coordinate ground truth sequence to obtain an augmented video sequence and an augmented coordinate ground truth sequence.
[0073] Step 5, use the augmented video sequence and the augmented coordinate ground truth sequence to train the hierarchical inter-frame sliding window causal attention network model.
[0074] Step 6, use the trained hierarchical inter-frame sliding window causal attention network model for autoregressive target tracking.
[0075] Among them, in step 1, the video sampling frequency of the data set is set to 10 frames per second.
[0076] Among them, in step 2, when the occlusion ratio is greater than 50% and there are at least frames from the end frame of the video, it is determined as a valid frame, is a natural number greater than or equal to 48; when randomly selecting a valid frame from all valid frames as the reference frame, sampling optimization is performed based on the variance of the occlusion ratio within the interval,
[0077] Specifically as follows, set the window length to For each frame, traverse the occlusion ratio of all frames, calculate the variance within its corresponding window, filter out the variance calculation weights at the positions of valid frames, and normalize the weights to form a probability distribution:
[0078] ,
[0079] where, represents the probability that the th valid starting frame is selected, represents the variance of the occlusion ratio of the window corresponding to the th valid starting frame, represents the variance of the occlusion ratio of the window corresponding to the th valid starting frame;
[0080] Calculate the probability distribution, and use the random weighted sampling method to select the final reference frame;
[0081] The initial video sequence is defined as , where, represents the th frame starting from the reference frame, and the initial coordinate ground truth sequence , where, represents the initial coordinate ground truth of the th frame.
[0082] Among them, in step 3: Coordinate perturbation uses the dynamic noise injection technique, and the initial coordinate ground truth of the th frame in the sequence is expressed as:
[0083] ,
[0084] where, is the abscissa of the upper left vertex of the th frame box, is the ordinate of the upper left vertex of the th frame box, is the width of the th frame box, is the height of the th frame, and the initial coordinate ground truth of the th frame after perturbation is calculated as:
[0085] ,
[0086] ,
[0087] ,
[0088] ,
[0089] ,
[0090] Among them, is the abscissa perturbation of the upper-left vertex of the th frame box, is the ordinate perturbation of the upper-left vertex of the th frame box, is the perturbation width of the th frame box, is the perturbation height of the th frame; represents the natural constant, represents the sampling value of Gaussian noise with a variance of 1, represents the sampling value of uniform noise with a value range between 0 and 1;
[0091] The search area cropping adopts an adaptive magnification strategy. Centered on the perturbed coordinates, the cropping side length of the th frame is calculated as:
[0092] ,
[0093] Among them, is the cropping factor for controlling the cropping scale;
[0094] Based on the cropping side length and perturbed coordinates, a cropped video sequence is generated, and its expression is:
[0095] ,
[0096] Among them, each element , is generated through the following:
[0097] ,
[0098] ,
[0099] represents performing a square cropping operation on the image with the center point of as the center coordinate and a side length of ;
[0100] The true value sequence of cropping coordinates is obtained through the following processing. When generating the cropped video sequence , the initial true value sequence of coordinates is subjected to equivalent mapping processing, and finally the true value sequence of cropping coordinates is formed.
[0101] This method addresses the issue of inconsistent training and inference scenarios in existing target tracking technologies by proposing an innovative technical improvement solution. Different from the traditional single-frame input mode, this method constructs a video sequence as the model input sample for the first time, and realizes the consistency optimization of training and inference through the following key technical breakthroughs: First, an innovative coordinate space perturbation strategy is introduced, and by applying random displacement deviations during the training phase, it effectively simulates the prediction coordinate errors generated during the actual inference process; Second, through the search area cropping module, the image preprocessing process during the inference phase is accurately reproduced, and an end-to-end training and validation closed-loop is constructed. This dual technical solution successfully bridges the feature distribution differences between the model training environment and the deployment environment, enables the model to have the ability of multi-frame parallel training, breaks through the limitations of the traditional autoregressive training mode in terms of temporal dependence, and improves the training efficiency by more than 300% compared with the frame-by-frame processing method, providing a new model optimization paradigm for the field of video target tracking.
[0102] Among them, in step 4: enhancing the video sequence is obtained by applying the same rotation angle, flipping direction, Gaussian blur, and color channel permutation spatial transformation to all frames of the cropped video sequence simultaneously;
[0103] Enhancing the coordinate ground truth sequence is obtained from the cropped coordinate ground truth sequence through the equivalent coordinate mapping when the cropped video sequence is transformed into the enhanced video sequence ;
[0104] Among them, in step 5, the hierarchical feature extraction network adopts a four-level progressive architecture. The first two levels use the local intra-frame window attention mechanism for fine-grained feature extraction, and the last two levels use the global inter-frame sliding window causal attention to achieve cross-frame information fusion;
[0105] Among them, the specific calculation steps of the global inter-frame sliding window causal attention module are as follows:
[0106] First, perform feature combination. For the feature input in the -th frame, it is calculated as
[0107] where
[0108] represents the feature of the -th frame picture, represents the feature of the coordinates of the -th frame, and represents concatenating the vectors along the first dimension;
[0109] Then, input the total feature sequence Calculated as:
[0110] ,
[0111] Inter-frame sliding window causal attention limit matrix The calculation process is as follows:
[0112] ,
[0113] ,
[0114] Among them, represents the attention of the th feature of the th frame input to the th feature of the th frame input. Specifically, represents the total number of sub-blocks of the feature, is a natural constant less than or equal to and greater than 0, which controls that the th frame feature can at most focus on the first frame features including itself;
[0115] Finally, the input total feature sequence , the inter-frame sliding window causal attention limit matrix , and the output feature of the attention module are calculated as:
[0116] ,
[0117] Among them, represents the transpose of the matrix, , and are all learnable parameters of the network model, represents the length of each block of features;
[0118] Inter-frame rotation position encoding is adopted, and intra-frame fixed position encoding is adopted; the prediction confidence uses the intersection over union of the predicted coordinates and the predicted ground truth as the ground truth to represent the confidence degree of the model for the current predicted coordinates, which is used to correct the tracking state during inference; the loss function design fuses the coordinate regression error and the prediction confidence error, and uses weighted summation to balance the contribution degrees of different loss terms.
[0119] Among them, in step 6, the incremental processing flow is adopted in the inference stage, and the most recent For the characteristic state of the frame, in the search area dynamic adjustment algorithm, the search area is initialized with the coordinates of the first frame. During subsequent predictions, the search area is updated by sliding average with the intersection-over-union ratio of the predicted coordinates of two adjacent frames as the weight. When the confidence level is lower than the threshold, the range of the search area is expanded.
[0120] The technical solution provided by the present invention will be further elaborated through specific embodiments below. The operating system used in our solution is Ubuntu 22.04.2 LTS, the deep learning development framework is PyTorch, and Python is used as the development language. The CPU used in the experiment is Intel Core i7-12800k, and the GPU is NVIDIA GeForce RTX 3090Ti 24G. During the training process, the algorithm runs in the python 3.11 environment, and the dataset parameters are as follows:
[0121] 。
[0122] As described above, the above are only the specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A single-object tracking method based on deep learning hierarchical video attention, characterized in that The method includes the following steps: Step 1: Read the dataset to obtain the original video data, the target coordinates in each frame, and the target occlusion ratio. Step 2: Use the target occlusion ratio to screen out the reference frames, and construct an initial video sequence with the reference frame as the first frame and an initial coordinate ground truth sequence. Step 3: Perturb the coordinates of the initial coordinate ground truth sequence, use the perturbed coordinates to crop the initial video sequence to obtain a cropped video sequence, and use the perturbed coordinates to transform the initial coordinate ground truth sequence to obtain a cropped coordinate ground truth sequence. Step 4: Perform data augmentation on the cropped video sequence and the cropped coordinate ground truth sequence to obtain an augmented video sequence and an augmented coordinate ground truth sequence. Step 5: Use the augmented video sequence and the augmented coordinate ground truth sequence to train a hierarchical inter-frame sliding window causal attention network model. Step 6: Use the trained hierarchical inter-frame sliding window causal attention network model for autoregressive target tracking. Among them, in Step 3, The coordinate perturbation adopts the dynamic noise injection technique, and the initial coordinate true value B of the i-th frame in the initial coordinate true value sequence T is i expressed as: B i = [x i , y i , w i , h i , where x i is the abscissa of the upper left vertex of the i-th frame box, y i is the ordinate of the upper left vertex of the i-th frame box, w i is the width of the i-th frame box, h i is the height of the i-th frame, and the true value of the initial coordinates of the perturbed i-th frame is calculated as: Among them, is the abscissa perturbation of the upper left vertex of the i-th frame box, is the ordinate perturbation of the upper left vertex of the i-th frame box, is the perturbation width of the i-th frame box, is the perturbation height of the i-th frame. e represents the natural constant, σ(1) represents the Gaussian noise sampling value with a variance of 1, and U(0, 1) represents the uniform noise sampling value with a value range between 0 and 1; The search area cropping adopts an adaptive magnification strategy. With the coordinates after perturbation as the center, the cropping side length sz of the i-th frame is i calculated as: where f is a cropping factor for controlling the cropping scale. Based on the cropping side length and the perturbed coordinates, a cropped video sequence Vc is generated, and its expression is: wherein, each element i ∈ [1, N] is generated by the following,[[]] Indicates performing a square cropping operation on the image Z with the center point as the center coordinates and the side length of sz; The cropped coordinate ground truth sequence Tc is obtained through the following processing. When generating the cropped video sequence Vc, perform an equivalent mapping process on the initial coordinate ground truth sequence T, and finally form the cropped coordinate ground truth sequence.
2. The single-object tracking method based on deep learning hierarchical video attention according to claim 1, characterized in that, In Step 1, the video sampling frequency of the dataset is set to 10 frames per second.
3. The single-object tracking method based on deep learning hierarchical video attention according to claim 1, characterized in that, In Step 2, when the occlusion ratio is greater than 50% and there are at least N frames away from the end frame of the video, it is determined as a valid frame, where N is a natural number greater than or equal to 48. When randomly selecting a valid frame from the valid frames as the reference frame, sampling optimization is performed based on the variance of the occlusion ratio within the interval. Specifically, the window length is set to N frames, the occlusion ratios of all frames are traversed, the variances within their corresponding windows are calculated, the variance calculation weights at the positions of the valid frames are screened out, and the weights are normalized to form a probability distribution: Among them, p i represents the probability that the i-th valid frame is selected, and var i represents the variance of the occlusion ratio of the window corresponding to the i-th valid frame, and var j represents the variance of the occlusion ratio of the window corresponding to the j-th valid frame; Calculate the probability distribution, and use the random weighted sampling method to select the final reference frame. The initial video sequence is defined as V = [Z1, Z2,..., Z i ,..., Z N , where Z i represents the i-th frame starting from the reference frame, and the initial coordinate ground truth sequence T = [B1, B2,..., B i ,..., B N , where B i represents the initial coordinate ground truth of the i-th frame.
4. The single-object tracking method based on deep learning hierarchical video attention according to claim 3, characterized in that, In Step 4, Enhanced video sequence V ct obtained by synchronously applying the same rotation angle, flipping direction, Gaussian blur, and color channel permutation spatial transformation to all frames of the cropped video sequence V c ; Enhanced coordinate true value sequence T ct From the cropped coordinate true value sequence T c After cropping the video sequence V c It is obtained after equivalent coordinate mapping when transformed into the enhanced video sequence V ct at that time.
5. The single-object tracking method based on deep learning hierarchical video attention according to claim 4, wherein In Step 5, for the hierarchical inter-frame sliding window causal attention network model, the first two levels use the local intra-frame window attention mechanism for fine-grained feature extraction, and the last two levels use the global inter-frame sliding window causal attention to achieve cross-frame information fusion. Among them, the specific calculation steps of the global inter-frame sliding window causal attention module are as follows. First, perform feature combination. For the feature F input in the i-th frame i , it is calculated as F i = cat(F i Z , F i B ) Among them, F i Z represents the feature of the i-th frame image, and F i B represents the feature of the i-th frame coordinates, and cat represents concatenating vectors along the first dimension; Then, the input total feature sequence X is calculated as: X = cat(F1, F2, ..., F i , ..., ..., F N ), The calculation process of the inter-frame sliding window causal attention restriction matrix M is as follows: Among them, represents the attention of the m N -th feature pair of the m P -th feature of the input of the m N -th frame to the n P -th feature of the input of the n -th frame. Specifically, P represents the total number of sub-blocks of the feature, and w is a natural constant that is less than or equal to N and greater than 0. Finally, input the total feature sequence X, the causal attention limit matrix M of the frame-interval sliding window, and the output feature X of the attention module out It is calculated as: Among them, (·) T represents the transpose of a matrix, and W q , W k and W v are all learnable parameters of the network model, and d represents the length of each piece of feature; Rotary position encoding is used between frames, and fixed position encoding is used within frames. The predicted confidence uses the intersection over union of the predicted coordinates and the predicted ground truth as the ground truth to represent the confidence degree of the model for the current predicted coordinates, and is used to correct the tracking state during inference. The loss function design fuses the coordinate regression error and the predicted confidence error, and uses the weighted summation method to balance the contribution degrees of different loss terms.
6. The single-object tracking method based on deep learning hierarchical video attention according to claim 5, characterized in that In step 6, during the inference stage, an incremental processing flow is adopted. The feature states of the most recent w frames are retained through a circular buffer mechanism. In the search area dynamic adjustment algorithm, the search area is initialized with the coordinates of the first frame. During subsequent predictions, the search area is updated by sliding average with the intersection over union (IoU) of the predicted coordinates of two adjacent frames as the weight. When the confidence level is lower than the threshold, the range of the search area is expanded.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the single-object tracking method based on deep learning hierarchical video attention described in any one of claims 1 to 6 above.
8. A computer-readable storage medium having computer instructions stored thereon, characterized in that, When the computer instruction is executed by the processor, it implements the single-object tracking method based on deep learning hierarchical video attention described in any one of claims 1 - 6.
Citation Information
Patent Citations
Scene text recognition method and device
CN114155527A
Occlusion-aware multi-object tracking
WO2022256150A1