A method for establishing a target-accurate tracking network in noisy environments

By using an end-to-end adaptive noise reduction tracking network and adversarial noise training, the accuracy and robustness issues of target trackers in noisy environments are solved, achieving accurate tracking in noisy environments and improving the tracker's performance in noisy environments.

CN116721127BActive Publication Date: 2025-12-02YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310530851.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-12
Publication Date
2025-12-02
Estimated Expiration
2043-05-12

AI Technical Summary

Technical Problem

Existing target trackers lack sufficient tracking accuracy and robustness in noisy environments, and cannot achieve accurate positioning and tracking in low-quality video.

Method used

Design an end-to-end adaptive noise reduction tracking network. Through adversarial noise training, a noise generator is used to add different types and levels of noise to the video image. The impact of noise is measured by BCE loss, and the tracker loss function is optimized to improve tracking accuracy and robustness in noisy environments.

Benefits of technology

While ensuring clean video tracking performance, it significantly improves tracking accuracy and robustness in noisy environments, increasing tracking success rate and accuracy by 3.7-5.5 percentage points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721127B_ABST
    Figure CN116721127B_ABST
Patent Text Reader

Abstract

This invention discloses a method for establishing a network for accurate target tracking in noisy environments, comprising the following steps: setting the number of training iterations for adversarial noise training, the type and magnitude of noise generated by the noise generator, and the balance coefficient of the tracker's loss function; feeding a pair of images into the noise generator, which adds noise to the image pairs according to the set parameters, and then feeding the processed image pairs into the backbone network for feature extraction; classifying and regressing the response map obtained after feature extraction to obtain the tracking result of the current frame, and then iteratively running this process until all video frames of all video sequences in the test dataset are traversed; recording and saving the tracking result of each frame, and quantitatively analyzing the tracking accuracy and tracking success rate. This invention, while maximizing the BCE loss value, minimizes the final tracker loss value through training, thereby enabling the tracker to handle tracking in various noisy environments while maintaining tracking performance on clean videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision, and in particular to a method for establishing a network for accurate target tracking in noisy environments. Background Technology

[0002] Target tracking is a hot topic in computer vision research, and real-time target tracking is widely used in various fields such as UAV reconnaissance, video surveillance, autonomous driving, and robotics. Target tracking is the process of finding targets of interest in a video frame sequence. Its main task is to use the target information given in the first frame to locate the target's position in subsequent frames using a tracking strategy. The difficulty of target tracking lies in the various challenges faced in practical applications, such as complex backgrounds, scale changes, occlusion and movement out of the field of view, rapid movement, parallel and non-parallel rotation, and target tracking in noisy environments. In practical applications, these interference factors affect the tracker's performance to varying degrees, leading to a decrease in accuracy and success rate.

[0003] As research deepens, more and more excellent tracking algorithms have been proposed to address various tracking challenges. However, the performance evaluation of target tracking rarely considers the degradation of video quality. Noise is a byproduct of video image acquisition, transmission, and processing. Many situations can cause noise in images, such as Gaussian noise caused by changes in lighting, magnetic field interference from electronic circuits, and high temperatures in sensors. Salt-and-pepper noise can also occur during video transmission, decoding, and reception due to AD conversion errors. Furthermore, harsh outdoor environments, such as rain, heavy snow, fog, and darkness, can also lead to varying degrees of degradation in tracking videos. For noisy video sequences, trackers often experience performance degradation because noise affects the accurate representation of target features, making accurate target localization and tracking impossible. Currently, some researchers have utilized noise to improve tracker performance. Fiza et al. used a regularization method based on input noise to reduce generalization error and overfitting. Similarly, Li et al. explored and designed a noise-aware framework to improve the tracker's recognition ability. Amirkhani et al. applied style transfer techniques combined with the original dataset to train the tracker, improving its robustness.

[0004] While some methods for improving tracker performance using noise have been proposed in recent years, performance testing of target trackers on noisy datasets remains lacking. Current mainstream target tracking frameworks can only achieve high tracking accuracy on clean datasets while maintaining real-time tracking capabilities; however, their accuracy drops significantly when tracking low-quality videos. Therefore, researching deep learning-based adversarial training methods for target tracking in noisy environments is crucial for improving the accuracy and robustness of trackers when facing low-quality videos. Summary of the Invention

[0005] The technical problem to be solved by this invention is: to design an end-to-end adaptive noise reduction tracking network for target tracking tasks in noisy environments, so that the network can achieve accurate target tracking in various noise environments.

[0006] To achieve the above technical objectives, the present invention adopts the following technical solution:

[0007] A method for establishing a network for accurate target tracking in noisy environments includes the following steps: First, setting the number of training iterations for adversarial noise training, the type and magnitude of noise generated by the noise generator, and the balance coefficient of the tracker's loss function; feeding a pair of images into the noise generator, which adds noise to the image pair according to the set parameters, and then feeding the processed image pair into the backbone network for feature extraction; classifying and regressing the response map obtained after feature extraction to obtain the tracking result of the current frame, and then iteratively running this process until all video frames of all video sequences in the test dataset are traversed; recording and saving the tracking result of each frame, and quantitatively analyzing the tracking accuracy and tracking success rate.

[0008] Furthermore, the method for establishing a target-accurate tracking network in a noisy environment includes the following steps:

[0009] Step 1: Acquire target tracking video. Send the clean video to the noise generator. The noise generator can add different types and levels of noise to the clean video image. After the noise is added, the generator will output the noise image and then read the next frame image for noise processing.

[0010] Step 2: Sample each video to ensure it contains clean data; the remaining data is then processed by adding noise using a noise generator.

[0011] Step 3: Assume that the xth frame of the video is image x, and this image has been processed by the noise generator. δ is the standard deviation of the noise, and x+δ represents the noisy image obtained after adding the standard deviation of the noise to the clean image of the xth frame.

[0012] Step 4: Input the bounding box of the template frame to detect targets in subsequent candidate regions. The template frame and the processed video frames to be detected are fed into the same backbone network. To maximize the obfuscation of the tracker by maximizing the noise generation effect of the video, a loss function, BCE loss, is introduced. BCE loss measures the impact of noise on the tracker and is defined as l. ce :

[0013] l ce = -ylg(p) + (1-y)lg(1-p)

[0014] Where y represents the classification score map of the frame to be detected after being extracted by the backbone network, and p represents the ground truth label;

[0015] Step 5: Let the bounding box result of the template frame be (x c y c w r h r ), where x c y c These are the x and y coordinates of the center point of the tracked image, respectively. r h r These are the width and height of the tracked image, respectively;

[0016] Step 6: Input the I-th frame of the video to be detected into the tracker to obtain N proposed candidate boxes. Calculate the N proposed candidate boxes of the current frame and the template frame result (x... c y c w r h r If the Intersection over Union (IOU) result is P, then the true classification confidence label is P. c :

[0017]

[0018] Step 7: For the nth (0 < n ≤ N) tracking proposal box of the current frame to be detected I in These are the x and y coordinates of the center point of the tracking result, respectively. These are the width and height of the tracking result, respectively, and their values ​​compared to the tracking result of the previous frame (x). c y c w r h r The true regression offset is Right now

[0019]

[0020]

[0021]

[0022]

[0023] Step 8: For the current frame to be detected I, the loss function of the nth (0 < n ≤ N) tracking proposal box obtained by the tracker is L(I, n, θ), that is;

[0024]

[0025] Among them, L cThis represents the binary classification loss function, calculated using the cross-entropy loss function; L r This represents the bounding box regression loss function, calculated using the smoothL1 loss function. This represents the prediction classification confidence score of the nth proposed candidate box in the current frame I to be detected, which represents the probability that the tracker predicts that the nth proposed candidate box in the current frame I to be detected contains the tracked target; The predicted regression offset of the nth suggestion candidate box in input frame I represents the offset between the tracker's predicted coordinates of the nth suggestion candidate box in input frame I and the target coordinates. This represents the true classification confidence score of the nth proposed candidate box in the current frame I, and represents the true probability that the nth proposed candidate box in the current frame I contains the tracked target; λ represents the true regression offset of the nth suggestion candidate box in the current frame I, which represents the true offset between the coordinates of the nth suggestion candidate box in the input frame I and the target coordinates. λ is a fixed weight parameter, which represents the network parameters used by the tracker.

[0026] Step 9: For the I-th frame image, calculate the highest predicted classification confidence score of the candidate boxes. The corresponding nth tracking suggestion box Add the corresponding predicted regression offset That is, the tracking result (x) pro y pro w pro h pro ), where the predicted regression bias Therefore, the tracking result is:

[0027]

[0028]

[0029]

[0030]

[0031] Step 10: Take the next frame I+1 of the video as the frame to be detected and repeat the operations of steps 4-9. Under the condition of maximizing the BCE loss value, minimize the final tracker loss value L(I, n, θ). Through training, continuously enable the tracker to cope with various noisy environments while ensuring the tracking effect on clean video.

[0032] Step 11: To prevent the noise generator from getting stuck in a local minimum, periodically stop the adversarial noise training (ANT), i.e., the operations from Step 1 to Step 10, and start training a new noise generator from scratch. The noise generator is trained based on the current state of the tracker to find the current optimal state. The noise generator is regarded as the inner loop and the tracker as the outer loop. The new noise generator replaces the previous noise generator in the adversarial noise training, which represents the update of the inner loop. After the inner loop is updated, the tracker of the outer loop will be updated accordingly. The two alternate until the network converges.

[0033] Furthermore, step 2 includes 50% clean data.

[0034] Furthermore, in step 2, the value of δ can be 0, 0.08, 0.12, 0.18, 0.26, 0.38, 0.5, 0.6 or 0.7.

[0035] Furthermore, the range of x+δ is [0,1].

[0036] The beneficial effects of this invention are:

[0037] (1) In the field of single-target tracking, most mainstream tracking frameworks can achieve good tracking results in clean videos, but the tracking accuracy will be greatly reduced when tracking targets in various noisy environments. How to achieve accurate tracking of targets in noisy environments is of great significance to improving the robustness of trackers in various scenarios. This invention focuses on single-target tracking in noisy environments and proposes an end-to-end adaptive noise reduction tracking network. Through adversarial noise training, the tracking accuracy and robustness of the tracker in noisy environments are improved without losing the tracking accuracy in clean videos.

[0038] (2) With the rapid development of computer vision, visual target tracking has been widely applied in many fields such as UAV reconnaissance, video surveillance, autonomous driving, and robotics. However, most mainstream tracking frameworks currently cannot achieve accurate target tracking in noisy environments. To address this issue, this invention proposes a target tracking algorithm based on adversarial noise training. The algorithm designs a noise generator that can add different types and levels of noise to clean video images. Furthermore, the algorithm measures the impact of noise on the tracker by introducing BCE loss between the classification score map and the ground truth label; the larger the loss value, the greater the impact of noise on the tracker. Under the condition of maximizing the BCE loss value, the final tracker loss value is minimized through training, thereby enabling the tracker to cope with tracking in various noisy environments while ensuring tracking performance on clean videos. Attached Figure Description

[0039] Figure 1Flowchart for an embodiment;

[0040] Figure 2 The success rate graph of the DaSiamRPN tracking algorithm in the embodiment before and after using noise adversarial training is shown.

[0041] Figure 3 The image shows the accuracy of the DaSiamRPN tracking algorithm in the example before and after using noise adversarial training. Detailed Implementation

[0042] The specific embodiments and working principles of the present invention will be further described in detail below with reference to the accompanying drawings.

[0043] Based on the above ideas, this embodiment provides a method for establishing a network for accurate target tracking in noisy environments, the workflow of which is as follows: Figure 1 As shown, the specific steps are as follows:

[0044] Step 1: Acquire the target tracking video. The clean video is then fed into a noise generator, which adds different types and levels of noise to the clean video image. Since the video is composed of several frames, the noise generator essentially operates on each frame. When a frame is fed into the noise generator, it adds different levels of Gaussian, uniform, salt-and-pepper, etc., noise to the clean image according to the required algorithm. After noise addition, the generator outputs the noisy image and then reads the next frame for further noise addition.

[0045] Step 2: To ensure high classification accuracy for clean samples, each video is sampled so that it contains 50% clean data, and the remaining 50% of the data is processed by adding noise using a noise generator.

[0046] Step 3: Assume that the x-th frame of the video is x, and this image has been processed by the noise generator. δ is the standard deviation of the noise. x+δ represents the noisy image obtained after adding the standard deviation of the noise to the clean image of the x-th frame. The possible values ​​of δ are 0, 0.08, 0.12, 0.18, 0.26, 0.38, 0.5, 0.6, and 0.7. The range of x+δ is [0,1].

[0047] Step 4: Input the bounding box of the template frame to detect targets in subsequent candidate regions. The template frame and the processed video frames to be detected are fed into the same backbone network. To maximize the obfuscation of the tracker by maximizing the noise generation, a loss function, BCE loss, is introduced. BCE loss measures the impact of added noise on the tracker. It is defined as l ce :

[0048] l ce = -ylg(p) + (1-y)lg(1-p)

[0049] Where y represents the classification score map of the frame to be detected after being extracted by the backbone network, and p represents the ground truth label. The magnitude of the loss value can be used to determine the extent of the noise's impact on the tracker. A larger loss value indicates that the noise generator is more likely to confuse the tracker; a smaller loss value indicates that the noise generator has a smaller impact on the tracker.

[0050] Step 5: Let the bounding box result of the template frame be (x c y c w r h r ), where x c y c These are the x and y coordinates of the center point of the tracked image, respectively. r h r These represent the width and height of the tracking image, respectively.

[0051] Step 6: Input the I-th frame of the video to be detected into the tracker to obtain N proposed candidate boxes. Calculate the N proposed candidate boxes of the current frame and the template frame result (x... c y c w r h r If the Intersection over Union (IOU) result is P, then the true classification confidence label is P. c :

[0052]

[0053] Step 7: For the nth (0 < n ≤ N) tracking proposal box of the current frame to be detected I in These are the x and y coordinates of the center point of the tracking result, respectively. These represent the width and height of the tracking result, respectively. This is compared to the tracking result of the previous frame (x...). c y c w r h r The true regression offset is Right now

[0054]

[0055]

[0056]

[0057]

[0058] Step 8: For the current frame to be detected I, the loss function of the nth (0 < n ≤ N) tracking proposal box obtained by the tracker is L(I, n, θ), that is;

[0059]

[0060] Among them, L c This represents the binary classification loss function, calculated using the cross-entropy loss function; L r This represents the bounding box regression loss function, calculated using the smoothL1 loss function. This represents the prediction classification confidence score of the nth proposed candidate box in the current frame I to be detected, which represents the probability that the tracker predicts that the nth proposed candidate box in the current frame I to be detected contains the tracked target; The predicted regression offset of the nth suggestion candidate box in input frame I represents the offset between the tracker's predicted coordinates of the nth suggestion candidate box in input frame I and the target coordinates. This represents the true classification confidence score of the nth proposed candidate box in the current frame I, and represents the true probability that the nth proposed candidate box in the current frame I contains the tracked target; This represents the true regression offset of the nth proposed candidate box in the current frame I, indicating the actual offset between the coordinates of the nth proposed candidate box in the input frame I and the target coordinates. λ is a fixed weight parameter, representing the network parameters used by the tracker.

[0061] Step 9: For the I-th frame image, calculate the highest predicted classification confidence score of the candidate boxes. The corresponding nth tracking suggestion box Add the corresponding predicted regression offset That is, the tracking result (x) pro y pro w pro h pro ), where the predicted regression bias Therefore, the tracking result is:

[0062]

[0063]

[0064]

[0065]

[0066] Step 10: Take the next frame I+1 of the video as the frame to be detected and repeat the operations of steps 4-9. Under the condition of maximizing the BCE loss value, minimize the final tracker loss value L(I, n, θ). Through training, continuously enable the tracker to cope with various noisy environments while ensuring the tracking effect on clean video.

[0067] Step 11: To prevent the noise generator from getting stuck in a local minimum, periodically stop the Adversarial Noise Training (ANT), i.e., the operations from Steps 1 to 10, and train a new noise generator from scratch. The noise generator is trained based on the current state of the tracker to find the current optimal state. Consider the noise generator as the inner loop and the tracker as the outer loop. The new noise generator replacing the previous noise generator in the ANT represents the update of the inner loop. After the inner loop is updated, the tracker in the outer loop will be updated accordingly. The two processes alternate until the network converges.

[0068] like Figure 2 and Figure 3 As shown, the legend “DaSiamRPN” represents the tracking performance of DaSiamRPN without training with adversarial noise, and the legend “DaSiamRPN_noise” represents the tracking performance of DaSiamRPN with training with adversarial noise. After training with adversarial noise, the tracking success rate and tracking accuracy of the DaSiamRPN algorithm increased by 3.7 and 5.5 percentage points, respectively.

Claims

1. A method for establishing a network for accurate target tracking in a noisy environment, characterized in that, Includes the following steps: First, set the number of training iterations for adversarial noise training, the type and magnitude of noise generated by the noise generator, and the balance coefficient of the tracker's loss function; then, feed a pair of images into the noise generator, which adds noise to the image pair according to the set parameters, and finally feed the processed image pair into the backbone network for feature extraction. The response map obtained after feature extraction is classified and regressed to obtain the tracking result of the current frame. This process is then iterated until all video frames of all video sequences in the test dataset are traversed. The tracking results for each frame are recorded and saved, and the tracking accuracy and tracking success rate are quantitatively analyzed. The method for establishing a target-accurate tracking network in a noisy environment includes the following steps: Step 1: Acquire target tracking video. Send the clean video to the noise generator. The noise generator can add different types and levels of noise to the clean video image. After the noise is added, the generator will output the noise image and then read the next frame image for noise processing. Step 2: Sample each video to ensure it contains clean data; the remaining data is then processed by adding noise using a noise generator. Step 3: Assume that the xth frame of the video is image x, and this image has been processed by the noise generator. δ is the standard deviation of the noise, and x+δ represents the noisy image obtained after adding the standard deviation of the noise to the clean image of the xth frame. Step 4: Input the bounding box of the template frame to detect targets in subsequent candidate regions; feed the template frame and the processed video frames to be detected into the same backbone network. To maximize the obfuscation of the tracker by processing the video with the noise generator, a loss function, BCE loss, is introduced. BCE loss measures the impact of noise on the tracker and is defined as l ce : l ce =-y lg(p)+(1-y)lg(1-p) Where y represents the classification score map of the frame to be detected after being extracted by the backbone network, and p represents the ground truth label; Step 5: Let the bounding box result of the template frame be (x c ,y c ,w r ,h r ), where x c , y c represents the x and y coordinates of the center point of the tracking image, and ω represents the x and y coordinates of the center point of the tracking image. r h r These are the width and height of the tracked image, respectively; Step 6: Input the I-th frame of the video to be detected into the tracker to obtain N proposed candidate boxes. Calculate the N proposed candidate boxes of the current frame and the template frame result (x... c y c w r h r If the Intersection over Union (IOU) result is P, then the true classification confidence label is P. c : Step 7: For the nth (0 < n ≤ N) tracking proposal box of the current frame to be detected I in These are the x and y coordinates of the center point of the tracking result, respectively. These are the width and height of the tracking result, respectively, and their values ​​compared to the tracking result of the previous frame (x). c y c w r h r The true regression offset is Right now Step 8: For the current frame to be detected I, the loss function of the nth (0 < n ≤ N) tracking proposal box obtained by the tracker is L(I, n, θ), that is; Among them, L c This represents the binary classification loss function, calculated using the cross-entropy loss function; L r This represents the bounding box regression loss function, calculated using the smoothL1 loss function. This represents the prediction classification confidence score of the nth suggested candidate box in the current frame I to be detected, which represents the probability that the tracker predicts that the nth suggested candidate box in the current frame I to be detected contains the tracked target; The predicted regression offset of the nth suggested candidate box in input frame I represents the offset between the tracker's predicted coordinates of the nth suggested candidate box in input frame I and the target coordinates. The true classification confidence score of the nth proposed candidate box in the current frame I represents the true probability that the nth proposed candidate box in the current frame I contains the tracked target; λ represents the true regression offset of the nth suggestion candidate box in the current frame I, and θ represents the true offset between the coordinates of the nth suggestion candidate box in the input frame I and the target coordinates. λ is a fixed weight parameter, and θ represents the network parameters used by the tracker. Step 9: For the I-th frame image, calculate the highest predicted classification confidence score of the candidate boxes. The corresponding nth tracking suggestion box Add the corresponding predicted regression offset That is, the tracking result (x) pro ,y pro ,w pro ,h pro ), where the predicted regression offset Therefore, the tracking result is: Step 10: Take the next frame I+1 of the video as the frame to be detected and repeat the operations of steps 4-9. Under the condition of maximizing the BCE loss value, minimize the final tracker loss value L(I, n, θ). Through training, continuously enable the tracker to cope with various noisy environments while ensuring the tracking effect on clean video. Step 11: To prevent the noise generator from getting stuck in a local minimum, periodically stop the adversarial noise training (ANT), i.e., the operations from Step 1 to Step 10, and start training a new noise generator from scratch. The noise generator is trained based on the current state of the tracker to find the current optimal state. The noise generator is regarded as the inner loop and the tracker as the outer loop. The new noise generator replaces the previous noise generator in the adversarial noise training, which represents the update of the inner loop. After the inner loop is updated, the tracker of the outer loop will be updated accordingly. The two alternate until the network converges.

2. The method for establishing a precise target tracking network in a noisy environment according to claim 1, characterized in that, Step 2 contains 50% clean data.

3. The method for establishing a precise target tracking network in a noisy environment according to claim 1, characterized in that, In step 2, the value of δ can be 0, 0.08, 0.12, 0.18, 0.26, 0.38, 0.5, 0.6 or 0.

7.

4. The method for establishing a precise target tracking network in a noisy environment according to claim 1, characterized in that, The range of x+δ is [0,1].

Citation Information

Patent Citations

  • Single target tracking method based on Siamese network

    CN111797716A

  • Video tracking-oriented attack resisting method and system, medium, equipment and terminal

    CN115511910A