An attention-guided adversarial patch generation method for secure detection in visual tracking
Through the sensitive detection module and patch attack module of the TrackSpear model, adversarial patches are accurately located and embedded, solving the problem of poor effectiveness of existing attack methods and improving the adversarial robustness and attack effect of the Transformer visual tracking model.
Patent Information
- Application Number
- CN202510176405.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing local perturbation attack methods are not effective for visual tracking models based on the Transformer structure. They cannot effectively perturb the specific attention area of the tracker and lack consideration of the self-attention mechanism of the Transformer model, which increases the difficulty of designing adversarial patches.
The TrackSpear model is adopted, which includes a sensitive detection module and a patch attack module. The key attack area is accurately located through the attention mechanism, and adversarial patches are generated and embedded. The Hadamard product and mask matrix are used to control the patch embedding range and optimize the tracking performance of the perturbation interference target tracker.
It significantly enhances the adversarial attack capability of Transformer-based target trackers, increases the threat level of the model and its sensitivity to potential attacks, and undermines the stability of the target tracking system.
Smart Images

Figure CN120164184B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving and machine vision target tracking, and relates to a method for generating adversarial patches for attention-guided visual tracking safety detection. Background Art
[0002] In autonomous driving traffic information systems, accurately tracking the movement of surrounding objects is crucial to ensuring the safety and reliability of autonomous vehicles. Visual object tracking, as a core task in computer vision, aims to locate and continuously track target objects in real time in dynamically changing video sequences, especially in complex and rapidly changing environments. With the rapid development of deep learning technology, Transformer-based object tracking models can capture long-range dependencies between the target and the surrounding environment through the self-attention mechanism, showing excellent performance, especially when dealing with dynamic and complex scenes. However, deep neural networks, especially visual tracking models, are vulnerable to attacks using carefully designed adversarial samples. With the emergence of physical adversarial attacks, this threat has become more feasible in real-world scenarios.
[0003] Existing adversarial attack methods mostly rely on global perturbations, but in practical applications, global perturbations are difficult to implement because the attacks require a high degree of physical feasibility and accuracy. In practice, attacks usually rely on local patches to disrupt target tracking. However, existing local perturbation attack methods are less effective against visual tracking models based on the Transformer structure. Because the Transformer model can capture global dependencies and has strong adversarial robustness, it is usually able to resist small-scale perturbations, making it more difficult to design effective physical adversarial patches. Therefore, attack methods targeting the Transformer structure not only help to gain a deeper understanding of its potential weaknesses, but also provide an important research direction for further improving its defense capabilities. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this paper proposes an attention-guided adversarial patch generation method for visual tracking security detection. This method aims to address the problem that existing local perturbation attack methods are ineffective against Transformer-based visual tracking models.
[0005] Specifically, the technical issues include: the threat of adversarial samples: With the popularization of autonomous driving technology, visual target tracking systems face the challenge of adversarial samples. Carefully designed adversarial samples can deceive tracking models through local perturbations in real-world scenarios, resulting in errors in target tracking, which in turn affects the safety and reliability of autonomous driving systems; the lack of consideration of the attention distribution and target characteristics of target trackers: different trackers focus on different locations, and existing adversarial patches do not fully consider this, resulting in limited attack effectiveness and inability to effectively perturb the specific attention area of the tracker; lack of consideration for the self-attention mechanism of the Transformer model: the Transformer model captures global dependencies through the self-attention mechanism and has strong adversarial robustness. Existing research has not fully considered how to design effective adversarial patches based on this feature.
[0006] The technical solutions of the present invention are as follows:
[0007] An attention-guided adversarial patch generation method for visual tracking security detection is implemented through the TrackSpear model, which consists of two main modules: a sensitive detection module and a patch attack module.
[0008] The sensitive detection module detects the sensitive positions of the target area in the video frame through the attention mechanism and accurately locates the key attack area;
[0009] The patch attack module generates and embeds adversarial patches to interfere with the tracking performance of the target tracker by optimizing the perturbations;
[0010] The specific steps are as follows:
[0011] (1) Video sequence I = {I1, I2, ..., I t Input TrackSpear model for processing and analysis;
[0012] (2) The sensitive detection module analyzes the video frame by frame and generates the attention map A based on the tracker structure. t , accurately locate key attack areas p * ;
[0013] (3) In the patch generation module, the perturbation generator G is based on the video frame I t Generate a specific adversarial patch p = G(I t ), and embed it into the key attack area p in step (2) * , generate adversarial samples Where ⊙ is the Hadamard product and m is the mask matrix, which is used to control the embedding range of the patch to achieve effective attack on target tracking.
[0014] Preferably, the specific steps of the above step (2) are as follows:
[0015] Depending on the tracker structure, the generation of attention maps and the location of key attack areas are divided into the following two methods:
[0016] (2-1) For the corner nodding method: analyze the attention map and select the area with the highest attention. Based on the center point of the template area as a reference, capture the relationship between different tokens in the search area and generate an attention map from the search area to the template area. Where: Q represents the query vector in the search area, k represents the key vector in the template area, d k is the dimension of the key vector, which is used as a scaling factor to ensure the stability of the calculation. According to the calculated attention map, the patch placement position p is calculated. * , ensuring that the patch can be accurately embedded in the key target position in the search area;
[0017] (2-2) For the central head method: generating an attention map Select the position with the highest attention value as the patch placement point p * ,Since the attention map in the center head method directly reflects the center position of the target object, there is no need to refer to the center point of the template area, where,V,is a value vector containing the feature information associated with each position, which is used to weightedly generate the final feature representation.
[0018] Preferably, the specific training process of the disturbance generator G in step (3) above is as follows:
[0019] (3-1) Dataset Construction: Collect diverse visual datasets, including publicly available dashcam datasets, autonomous driving datasets, and actual vehicle driving video data, to ensure the broad adaptability and good generalization ability of the generated model, thereby improving the performance and reliability of the model in complex scenarios;
[0020] (3-2) Model sequence input and initialization: Extract video frame sequences {X1, X2, …, X t}, as the input of the generator G and its parameters φ during training;
[0021] (3-3) Generate adversarial patches: For each frame X t , generate the adversarial patch p through the perturbation generator G t , the specific formula is:
[0022] p t =G(I t ;φ)
[0023] Next, the generated adversarial patch p t Embedded in frame Xt Key attack areas, generating adversarial samples The formula is: Where ⊙ is the Hadamard product and m is the mask matrix used to control the embedding range of the patch;
[0024] (3-4) Calculate the loss function: define the total loss function as L = αL Prod +βL Cl s+γL Reg , where L Prod is the dot product loss for the attention mechanism, L Cls is the classification loss, L Reg For regression loss, coefficients α, β, and γ are used to balance the impact of each loss term;
[0025] (3-5) Parameter update: Use the Adam optimizer to update the parameter φ of the perturbation generator G. The specific formula is:
[0026]
[0027] Where η is the learning rate;
[0028] Repeat (3-2) to (3-5) until the generator G converges or reaches the maximum number of training iterations.
[0029] Preferably, the dot product loss L for the attention mechanism is Prod The specific algorithm is as follows:
[0030]
[0031] Matrix A h Represents the attention matrix of the self-attention layer l and the attention head h, the query matrix Q and the key matrix K represent the query and keyword vectors respectively; to prevent large dot product values from causing gradient explosion or disappearance, Q and K use ∥.∥ 1,2 The norm is normalized to ensure gradient stability. Here n represents the sequence length, and the matrix X refers to the vector, which can be Q or K. L layers Represents the total number of layers in the self-attention mechanism, H heads Indicates the number of attention heads in each layer;
[0032] Before being input into the Transformer model, the image is divided into fixed-size patches. Each patch is embedded into a fixed-dimensional vector space through a linear mapping function f. The mapping between patches and tokens follows row-major order, traversing from left to right and top to bottom, thereby determining the token index based on the patch position.
[0033] In the target tracking task, the attack can be implemented through three key positions. After the patch is added to the search area, the model first performs self-attention calculation and then applies cross-attention. In the self-attention layer, the attack can be launched from the query matrix Q or the key matrix K. When attacking from the query side, the model will focus more attention on the patch position, amplifying the effect of the patch on the target features and disrupting target detection. When attacking from the key side, the key vector is disturbed, affecting the key-value mapping, amplifying the patch's attraction to other areas, and changing the self-attention distribution.
[0034] In the cross-attention layer, the attack improves the similarity between the patch and the template region by misleading the model, mistakenly identifying the patch as the target. Prod The loss function further enhances the adversarial effect by implementing attacks on the key side of the search area.
[0035] Preferably, the classification loss L Cls The specific algorithm is as follows:
[0036]
[0037] in, is the probability feature map generated by the tracker from the original samples of frame t, and denote the probability feature map and classification feature map generated by the adversarial sample transfer tracker at frame t, respectively. H denotes the input feature region whose confidence exceeds the threshold δ. Here, δ is the confidence threshold, λ is the weight coefficient, and Q is an additional constraint term introduced to adjust the performance of the overall loss function.
[0038] Loss function L Cls Binary cross entropy is used to measure the area with confidence higher than δ, i.e. The difference between the foreground and background scores in high-confidence regions is minimized by adding the constraint term Q, which forces the model to make it more difficult to distinguish between foreground and background, thereby interfering with the model decision and increasing its vulnerability to adversarial perturbations.
[0039] Preferably, the regression loss L Reg The specific algorithm is as follows:
[0040]
[0041] and Represents the regression feature map generated by the original sample and the adversarial sample transfer tracker at frame t, bbox gt and,bbox pred[H] represents the predicted bounding boxes generated by the adversarial sample and the original sample transmission tracker m, respectively. During the tracking process, a low IoU value between the predicted box and the ground truth box usually indicates that the predicted box is not suitable as the final tracking result. Compared with IoU, GIoU provides a more significant improvement. Even if the predicted box is completely off the target, GIoU can effectively measure the offset between the predicted box and the true target. As the relative distance between the two increases, the GIoU value gradually increases. This feature helps guide the tracker's prediction away from the true target position;
[0042] In order to interrupt the tracking process, first select the bbox whose confidence is greater than the threshold δ gt The bounding box in the area and use bbox pred [H] Calculates the GIoU value of the actual target location; this process causes the selected prediction box to deviate from the true target location and reduce its width and height. As a result, in the next frame, the search area may no longer contain the location of the true target, thereby weakening the performance of the tracker.
[0043] Beneficial effects of the present invention:
[0044] This paper introduces an attention-based perception strategy and an attention loss function to effectively enhance the ability of the Transformer-based target tracker to resist attacks, thereby accurately destroying the stability of the target tracking system. In addition, this paper can significantly improve the threat of the Transformer-based visual tracking model and enhance its sensitivity to potential attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is the algorithm flow chart of the present invention. DETAILED DESCRIPTION
[0046] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.
[0047] like Figure 1 As shown in the figure, an attention-guided adversarial patch generation method for visual tracking security detection is implemented through the TrackSpear model. The model contains two main modules: a sensitive detection module and a patch attack module.
[0048] The sensitive detection module detects the sensitive positions of the target area in the video frame through the attention mechanism and accurately locates the key attack area;
[0049] The patch attack module generates and embeds adversarial patches to interfere with the tracking performance of the target tracker by optimizing the perturbations;
[0050] The specific steps are as follows:
[0051] (1) Video sequence I = {I1, I2, ..., I t Input TrackSpear model for processing and analysis;
[0052] (2) The sensitive detection module analyzes the video frame by frame and generates the attention map A based on the tracker structure. t , accurately locate key attack areas p * ;
[0053] (3) In the patch generation module, the perturbation generator G is based on the video frame I t Generate a specific adversarial patch p = G(I t ), and embed it into the key attack area p in step (2) * , generate adversarial samples Where ⊙ is the Hadamard product and m is the mask matrix, which is used to control the embedding range of the patch to achieve effective attack on target tracking.
[0054] Preferably, the specific steps of the above step (2) are as follows:
[0055] Depending on the tracker structure, the generation of attention maps and the location of key attack areas are divided into the following two methods:
[0056] (2-1) For the corner nodding method: analyze the attention map and select the area with the highest attention. Based on the center point of the template area as a reference, capture the relationship between different tokens in the search area and generate an attention map from the search area to the template area. Where: Q represents the query vector in the search area, k represents the key vector in the template area, d k is the dimension of the key vector, which is used as a scaling factor to ensure the stability of the calculation. According to the calculated attention map, the patch placement position p is calculated. * , ensuring that the patch can be accurately embedded in the key target position in the search area;
[0057] (2-2) For the central head method: generating an attention map Select the position with the highest attention value as the patch placement point p * ,Since the attention map in the center head method directly reflects the center position of the target object, there is no need to refer to the center point of the template area, where,V,is a value vector containing the feature information associated with each position, which is used to weightedly generate the final feature representation.
[0058] Preferably, the specific training process of the disturbance generator G in step (3) above is as follows:
[0059] (3-1) Dataset Construction: Collect diverse visual datasets, including publicly available dashcam datasets, autonomous driving datasets, and actual vehicle driving video data, to ensure the broad adaptability and good generalization ability of the generated model, thereby improving the performance and reliability of the model in complex scenarios;
[0060] (3-2) Model sequence input and initialization: Extract video frame sequences {X1, X2, …, X t}, as the input of the generator G and its parameters φ during training;
[0061] (3-3) Generate adversarial patches: For each frame X t , generate the adversarial patch p through the perturbation generator G t , the specific formula is:
[0062] p t =G(I t ;φ)
[0063] Next, the generated adversarial patch p t Embedded in frame X t Key attack areas, generating adversarial samples The formula is: Where ⊙ is the Hadamard product and m is the mask matrix used to control the embedding range of the patch;
[0064] (3-4) Calculate the loss function: define the total loss function as L = αL Prod +βL Cls +γL Reg , where L Prod is the dot product loss for the attention mechanism, L Cls is the classification loss, L Reg For regression loss, coefficients α, β, and γ are used to balance the impact of each loss term;
[0065] (3-5) Parameter update: Use the Adam optimizer to update the parameter φ of the perturbation generator G. The specific formula is:
[0066]
[0067] Where η is the learning rate;
[0068] Repeat (3-2) to (3-5) until the generator G converges or reaches the maximum number of training iterations.
[0069] Preferably, the dot product loss L for the attention mechanism is Prod The specific algorithm is as follows:
[0070]
[0071] Matrix A h Represents the attention matrix of the self-attention layer l and the attention head h, the query matrix Q and the key matrix K represent the query and keyword vectors respectively; to prevent large dot product values from causing gradient explosion or disappearance, Q and K use ∥.∥ 1,2 The norm is normalized to ensure gradient stability. Here n represents the sequence length, and the matrix X refers to the vector, which can be Q or K. L layers Represents the total number of layers in the self-attention mechanism, H heads Indicates the number of attention heads in each layer;
[0072] Before being input into the Transformer model, the image is divided into fixed-size patches. Each patch is embedded into a fixed-dimensional vector space through a linear mapping function f. The mapping between patches and tokens follows row-major order, traversing from left to right and top to bottom, thereby determining the token index based on the patch position.
[0073] In the target tracking task, the attack can be implemented through three key positions. After the patch is added to the search area, the model first performs self-attention calculation and then applies cross-attention. In the self-attention layer, the attack can be launched from the query matrix Q or the key matrix K. When attacking from the query side, the model will focus more attention on the patch position, amplifying the effect of the patch on the target features and disrupting target detection. When attacking from the key side, the key vector is disturbed, affecting the key-value mapping, amplifying the patch's attraction to other areas, and changing the self-attention distribution.
[0074] In the cross-attention layer, the attack improves the similarity between the patch and the template region by misleading the model, mistakenly identifying the patch as the target. Prod The loss function further enhances the adversarial effect by implementing attacks on the key side of the search area.
[0075] Preferably, the above classification loss L Cls The specific algorithm is as follows:
[0076]
[0077] in, is the probability feature map generated by the tracker from the original samples of frame t, and denote the probability feature map and classification feature map generated by the adversarial sample transfer tracker at frame t, respectively. H denotes the input feature region whose confidence exceeds the threshold δ. Here, δ is the confidence threshold, λ is the weight coefficient, and Q is an additional constraint term introduced to adjust the performance of the overall loss function.
[0078] Loss function L ClsBinary cross entropy is used to measure the area with confidence higher than δ, i.e. The difference between the foreground and background scores in high-confidence regions is minimized by adding the constraint term Q, which forces the model to make it more difficult to distinguish between foreground and background, thereby interfering with the model decision and increasing its vulnerability to adversarial perturbations.
[0079] Preferably, the regression loss L Reg The specific algorithm is as follows:
[0080]
[0081] and Represents the regression feature map generated by the original sample and the adversarial sample transfer tracker at frame t, bbox ge and,bbox pred [H] represents the predicted bounding boxes generated by the adversarial sample and the original sample transmission tracker m, respectively. During the tracking process, a low IoU value between the predicted box and the ground truth box usually indicates that the predicted box is not suitable as the final tracking result. Compared with IoU, GIoU provides a more significant improvement. Even if the predicted box is completely off the target, GIoU can effectively measure the offset between the predicted box and the true target. As the relative distance between the two increases, the GIoU value gradually increases. This feature helps guide the tracker's prediction away from the true target position;
[0082] In order to interrupt the tracking process, first select the bbox whose confidence is greater than the threshold δ gt The bounding box in the area and use bbox pred [H] Calculates the GIoU value of the actual target location; this process causes the selected prediction box to deviate from the true target location and reduce its width and height. As a result, in the next frame, the search area may no longer contain the location of the true target, thereby weakening the performance of the tracker.
[0083] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for generating adversarial patches for attention-guided visual tracking security detection, characterized by: Adversarial patch generation is achieved through the TrackSpear model, which consists of two main modules: a sensitive detection module and a patch attack module; The sensitive detection module detects the sensitive positions of the target area in the video frame through the attention mechanism and accurately locates the key attack area; The patch attack module generates and embeds adversarial patches to interfere with the tracking performance of the target tracker by optimizing the perturbations; The specific steps are as follows: (1) Video sequence I = {I1, I2, ..., I t Input TrackSpear model for processing and analysis; (2) The sensitive detection module analyzes the video frame by frame and generates the attention map A based on the tracker structure. t , accurately locate key attack areas p * ; (3) In the patch generation module, the perturbation generator G is based on the video frame I t Generate a specific adversarial patch p = G(I t ), and embed it into the key attack area p in step (2) * , generate adversarial samples Where ⊙ is the Hadamard product and m is the mask matrix, which is used to control the embedding range of the patch to achieve effective attack on target tracking; The specific steps of step (2) are as follows: Depending on the tracker structure, the generation of attention maps and the location of key attack areas are divided into the following two methods: (2-1) For the corner nodding method: analyze the attention map and select the area with the highest attention. Based on the center point of the template area as a reference, capture the relationship between different tokens in the search area and generate an attention map from the search area to the template area. Where: Q represents the query vector in the search area, k represents the key vector in the template area, d k is the dimension of the key vector, which is used as a scaling factor to ensure the stability of the calculation. According to the calculated attention map, the patch placement position p is calculated. * , ensuring that the patch can be accurately embedded in the key target position in the search area; (2-2) For the central head method: generating an attention map Select the position with the highest attention value as the patch placement point p * ,Since the attention map in the center head method directly reflects the center position of the target object, there is no need to refer to the center point of the template area, where,V,is a value vector containing the feature information associated with each position, which is used to weightedly generate the final feature representation.
2. The method for generating adversarial patches for attention-guided visual tracking security detection according to claim 1, characterized in that: The specific training process of the disturbance generator G in step (3) is as follows: (3-1) Dataset Construction: Collect diverse visual datasets, including publicly available dashcam datasets, autonomous driving datasets, and actual vehicle driving video data, to ensure the broad adaptability and good generalization ability of the generated model, thereby improving the performance and reliability of the model in complex scenarios; (3-2) Model sequence input and initialization: Extract video frame sequences {X1, X2, …, X t }, as the input of the generator G and its parameters φ during training; (3-3) Generate adversarial patches: For each frame X t , generate the adversarial patch p through the perturbation generator G t , the specific formula is: p t =G(I t ;φ) Next, the generated adversarial patch p t Embedded in frame X t Key attack areas, generating adversarial samples The formula is: Where ⊙ is the Hadamard product and m is the mask matrix used to control the embedding range of the patch; (3-4) Calculate the loss function: define the total loss function as L = αL Prod +βL Cls +γL Reg , where L Prod is the dot product loss for the attention mechanism, L Cls is the classification loss, L Reg For regression loss, coefficients α, β, and γ are used to balance the impact of each loss term; (3-5) Parameter update: Use the Adam optimizer to update the parameter φ of the perturbation generator G. The specific formula is: Where η is the learning rate; Repeat (3-2) to (3-5) until the generator G converges or reaches the maximum number of training iterations.
3. The method for generating adversarial patches for attention-guided visual tracking security detection according to claim 2, characterized in that: Dot product loss L for attention mechanism Prod The specific algorithm is as follows: Matrix A h Represents the attention matrix of the self-attention layer l and the attention head h, the query matrix Q and the key matrix K represent the query and keyword vectors respectively; to prevent large dot product values from causing gradient explosion or disappearance, Q and K use ||.|| 1,2 The norm is normalized to ensure gradient stability; here n represents the sequence length, and the matrix X refers to the vector, which can be Q or K; L layers Represents the total number of layers in the self-attention mechanism, H heads Indicates the number of attention heads in each layer; Before being input into the Transformer model, the image is divided into fixed-size patches. Each patch is embedded into a fixed-dimensional vector space through a linear mapping function f. The mapping between patches and tokens follows row-major order, traversing from left to right and top to bottom, thereby determining the token index based on the patch position. In the target tracking task, the attack can be implemented through three key positions. After the patch is added to the search area, the model first performs self-attention calculation and then applies cross-attention. In the self-attention layer, the attack can be launched from the query matrix Q or the key matrix K. When attacking from the query side, the model will focus more attention on the patch position, amplifying the effect of the patch on the target features and disrupting the target detection. When attacking from the key side, the key vector is disturbed, affecting the key-value mapping, amplifying the patch's attraction to other areas, and changing the self-attention distribution. In the cross-attention layer, the attack improves the similarity between the patch and the template region by misleading the model, mistakenly identifying the patch as the target. Prod The loss function further enhances the adversarial effect by implementing attacks on the key side of the search area.
4. The method for generating adversarial patches for attention-guided visual tracking security detection according to claim 2, characterized in that: Classification loss L Cls The specific algorithm is as follows: in, is the probability feature map generated by the tracker from the original samples of frame t, and denote the probability feature map and classification feature map generated by the adversarial sample transfer tracker at frame t, respectively. H denotes the input feature region whose confidence exceeds the threshold δ. Here, δ is the confidence threshold, λ is the weight coefficient, and Q is an additional constraint term introduced to adjust the performance of the overall loss function. Loss function L Cls Binary cross entropy is used to measure the area with confidence higher than δ, i.e. The difference between the foreground and background scores in high-confidence regions is minimized by adding the constraint term Q, which forces the model to make it more difficult to distinguish between foreground and background, thereby interfering with the model decision and increasing its vulnerability to adversarial perturbations.
5. The method for generating adversarial patches for attention-guided visual tracking security detection according to claim 2, characterized in that: Regression loss L Reg The specific algorithm is as follows: and Represents the regression feature map generated by the original sample and the adversarial sample transfer tracker at frame t, bbox gt and,bbox pred [H] represents the predicted bounding boxes generated by the adversarial sample and the original sample transmission tracker m, respectively. During the tracking process, a low IoU value between the predicted box and the ground truth box usually indicates that the predicted box is not suitable as the final tracking result. Compared with IoU, GIoU provides a more significant improvement; even if the predicted box is completely off the target, GIoU can effectively measure the offset between the predicted box and the true target. As the relative distance between the two increases, the GIoU value gradually increases. This feature helps guide the tracker's prediction away from the true target position; In order to interrupt the tracking process, first select the bbox whose confidence is greater than the threshold δ gt The bounding box in the area and use bbox pred [H] Calculates the GIoU value of the actual target location; this process causes the selected prediction box to deviate from the true target location and reduce its width and height. As a result, in the next frame, the search area may no longer contain the location of the true target, thereby weakening the performance of the tracker.
Citation Information
Patent Citations
Anti-patch concealment enhancement method based on thermodynamic diagram and style migration
CN115995035A
Sample attack resisting method, system and terminal for visual target tracking
CN117274769A