A Siamese Region Proposal Network Model for UAV Target Tracking

By adding stripe pooling module and global context network module to the SiamRPN network, the stability and reliability problems in drone target tracking are solved, especially in complex scenarios, and the tracking accuracy and success rate are significantly improved.

CN114266805BActive Publication Date: 2025-05-13SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111664859.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-13
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In complex scenarios, there are stability and reliability problems in drone target tracking, especially in the case of light changes, occlusion and background interference, target drift is more serious.

Method used

Based on the SiamRPN network, a stripe pooling module and a global context network module are added to enhance the network's utilization of spatial information and the establishment of remote context relationships, thereby improving the accuracy and success rate of target tracking.

Benefits of technology

Through the improved network structure, the target drift problem is effectively alleviated, the accuracy and success rate of drone target tracking is improved, and the performance is excellent in the case of large background interference and lighting changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266805B_ABST
    Figure CN114266805B_ABST
Patent Text Reader

Abstract

The present invention discloses a twin region proposal network model for unmanned aerial vehicle target tracking. The network model adds a strip pooling module and a global context network module on the basis of a SiamRPN network, so that the network solves the remote dependency problem and effectively understands different tracking scenarios. Then, the calculation method of the intersection-over-union ratio is optimized to complete the feature extraction of the target and regress an accurate prediction frame. The present invention has strong robustness under the conditions of illumination changes, background interference and rapid target movement. When tested on the UAV123 public data set benchmark, the tracking speed is about 106 frames per second, and an accuracy rate of 0.754 and a success rate of 0.542 are obtained. Especially in the background interference environment, the accuracy rate and the success rate are increased by 8.29% and 11.63% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicle target tracking, and in particular relates to a twin region proposal network model for unmanned aerial vehicle target tracking. Background Art

[0002] In the era of intelligence, drones are widely used in military, unmanned driving, aerial photography, traffic monitoring, pesticide spraying, target following, human-computer interaction and autonomous driving. Drone target tracking is based on video images to screen and locate the area of ​​interest. In complex scenes, due to the influence of lighting, occlusion and the rapid movement of small targets, how to meet the stability and reliability of drone image tracking is an important research direction at present.

[0003] The purpose of visual tracking is to accurately estimate the position of the target object in the subsequent frames of the video image based on the bounding box given by the first frame of the current video image. The target tracking algorithm based on correlation filtering originated from the MOSS algorithm, which first introduced correlation filtering into the target tracking algorithm. The CSK algorithm introduced the kernel circulant matrix and determined the similarity between two adjacent frames by calculating the Gaussian kernel correlation matrix to achieve target tracking. The KCF algorithm introduced the kernel technique and multi-channel feature processing to track the target, which greatly simplified the amount of calculation in the tracking process and laid the theoretical and practical foundation for the subsequent correlation filter target tracking algorithm. The Alexnet network proposed in 2012 is a milestone in the development of deep learning. In deep learning, the correlation target tracking algorithm represented by SiamFC can achieve a good balance between accuracy and speed. It adopts a fully convolutional neural network structure and measures the similarity by matching the template frame with the test frame to locate the target. SiamRPN solves the original multi-scale problem by adding the RPN (region proposal network) network on the basis of SiamFC; however, it does not consider the use of spatial information by the network itself. Therefore, when the target has problems such as illumination changes, background interference and occlusion, the target will drift. Summary of the invention

[0004] The purpose of the present invention is to provide a twin region proposal network model for UAV target tracking. The network model adds a strip pooling module and a global context network module on the basis of the SiamRPN network, so as to improve the accuracy and success rate of UAV target tracking.

[0005] To achieve the above purpose, the present invention specifically adopts the following technical solutions:

[0006] A twin region proposal network model for unmanned aerial vehicle target tracking, comprising a template branch unit and a search branch unit, wherein the template branch unit comprises a first convolution module, a strip pooling module, a second convolution module, a third convolution module, a first matching module and a first output module, and the search branch unit comprises a fourth convolution module, a global context network module, a fifth convolution module, a sixth convolution module, a second matching module and a second output module;

[0007] The first convolution module and the fourth convolution module constitute a twin network, the first convolution module is connected to the strip pooling module, and the fourth convolution module is connected to the global context network module;

[0008] The second convolution module, the third convolution module, the first matching module, the first output module, the fifth convolution module, the sixth convolution module, the second matching module and the second output module constitute a region proposal network, the second convolution module and the third convolution module are both connected to the first matching module, and the first matching module is connected to the first output module; the fifth convolution module and the sixth convolution module are both connected to the second matching module, and the second matching module is connected to the second output module; wherein the strip pooling module is respectively connected to the second convolution module and the fifth convolution module, the fourth convolution module is connected to the third convolution module, and the global context network module is connected to the sixth convolution module.

[0009] Furthermore, the classification loss function in the region proposal network adopts a cross entropy loss function, and the regression loss function is an L1 norm loss function.

[0010] Furthermore, the template branch unit input image size is 127×127×3, and the search branch unit input image size is 255×255×3.

[0011] Compared with the prior art, the present invention has the following beneficial effects:

[0012] (1) Adding the strip pooling module and the global context network module effectively establishes remote context relationships while reducing the amount of computation, expands the receptive field of the backbone network, and completes the foreground and background classification and bounding box regression of the region proposal network;

[0013] (2) By improving the calculation method of the intersection-over-union ratio, the problem of bounding box selection can be effectively alleviated during the training tracking stage. During the training process, an accurate intersection-over-union ratio calculation can be obtained, allowing the network to screen out accurate prediction boxes during the non-maximum suppression process.

[0014] Tested on the UAV123 public dataset benchmark, the tracking speed was about 106 frames per second, with an accuracy of 0.754 and a success rate of 0.542. Especially in the background interference environment, the accuracy and success rate were improved by 8.29% and 11.63% respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION

[0016] like Figure 1 As shown, a twin region proposal network model for drone target tracking provided in this embodiment includes a template branch unit and a search branch unit, the template branch unit includes a first convolution module, a strip pooling module, a second convolution module, a third convolution module, a first matching module and a first output module, the search branch unit includes a fourth convolution module, a global context network module, a fifth convolution module, a sixth convolution module, a second matching module and a second output module.

[0017] The first convolution module and the fourth convolution module constitute a twin network, which is used to input a template image and a test image, and then compare the two images to achieve target tracking. The first convolution module inputs a template image with a size of 127×127×3, and the fourth convolution module inputs a test image with a size of 255×255×3.

[0018] The second convolution module, the third convolution module, the first matching module, the first output module, the fifth convolution module, the sixth convolution module, the second matching module and the second output module constitute a region proposal network, and the region proposal network includes two subnetworks: (1) a network for classifying foreground and background, and (2) a network for regressing bounding boxes.

[0019] The template image input by the template branch unit passes through the first convolution module to output a feature map of size 6×6×256, and then the feature map is striped through the strip pooling module. Strip pooling is performed along the horizontal and vertical directions of the window. Through the long-to-narrow kernel, it is easy to establish a remote contextual relationship. Expanding the receptive field of the backbone network is beneficial to the classification of the target and background during tracking. Therefore, strip pooling can help the twin network capture contextual relationships during tracking, and then perform spatial dimension weighting for target feature extraction, so that the network automatically assigns a larger proportion of weight to the target position, enhances the network's ability to distinguish the target, and enables the network to further parse the tracking scene. The feature maps after strip pooling are input into the second convolution module and the fifth convolution module respectively.

[0020] The test image input by the search branch unit is processed by the fourth convolution module to output a 6×6×256 feature map, which is simultaneously input into the third convolution module and the global context network module. The global context network module can better establish the dependency relationship of the network remote context, and deepen the network's global understanding ability in the current drone tracking scene, automatically increase the proportion of channels related to the target features, and reduce the proportion of channels unrelated to the target features, change the dependency between different channels, and make the bounding box regression more accurate.

[0021] The feature maps output by the second convolution module and the third convolution module are input to the first matching module for matching and then output through the first output module; the feature maps output by the fifth convolution module and the sixth convolution module are input to the second matching module for matching and then output through the second output module.

[0022] The prediction of the bounding box directly affects the performance of video tracking. Intersection over Union is a commonly used indicator for target detection. It can not only score positive and negative samples, but also evaluate the distance between the output boundary prediction box and the target real boundary box. The calculation of the intersection over union can well reflect the effect of the prediction box and the real box in the tracking process, and evaluate the subsequent tracking indicators. This embodiment uses the distance intersection over union to calculate the bounding box, which can effectively alleviate the problem of training divergence of the intersection over union in target detection, and normalizes the smallest prediction box and the real boundary box to make the regressed bounding box more accurate. The classification loss function in the region proposal network adopts the cross-entropy loss function, and the regression loss function adopts the L1 norm loss function (smooth L1 loss).

[0023] When in use, the drone tracking steps are as follows:

[0024] (1) Load the network model of this embodiment (DAPsiamRPN), determine whether the network is the first frame image, extract the first frame image of the video with a size of 127×127×3 from the input image as the input of the template branch, and the search branch uses the image size of 255×255×3 as the input of the search branch.

[0025] (2) The input template branch image and the detection branch image are passed through the DAPSiamRPN network, and cross-correlation operations are performed in the classification branch and regression branch of the region proposal network to generate the final response k feature maps and 2k regressed bounding boxes, and the classification scores of the target and background are obtained. Through the regression of the bounding box, the size of the bounding box is optimized to obtain the position of the target.

[0026] (3) In the subsequent video images, the search area is expanded, and the feature map with the largest response to the previous frame of video image is found through the detection branch for subsequent tracking. If the tracking template needs to be updated, the above steps are repeated. Finally, it is determined whether it is the last frame of image. If so, the tracking ends.

[0027] Simulation experiment

[0028] The experimental platform is Ubuntu 16.04 LTS system, using pytorch version 1.4 as the deep learning framework, the device is Inter Core i7-9700F CPU 3.00GHz×8, and the single GPU is GeForce GTX 2060Super 8G.

[0029] The training data for this experiment is video data that meets the tracking scenario extracted from the ILSVRC2017_VID dataset and the Youtube-BB dataset. 44,976 video sequences are extracted from ILSVRC2017_VID and 904 video sequences are extracted from Youtube-BB. There are more than one million video images with real labels. During the training process, the Alexnet network is used as the pre-training model and as the backbone network for feature extraction of video images. Then 20 rounds of training are performed, each round has 12,000 iterations, and the total training time is 13 hours. Stochastic gradient descent uses stochastic gradient descent (SGD), and the momentum is set to 0.9. To prevent gradient explosion during training, the gradient clipping is set to 10, and the dynamic learning rate is set from 0.03 to 0.00001. The candidate boxes use five ratios of 0.33, 0.5, 1, 2, and 3 respectively. Only the first frame of the video is sent to the template branch for template collection. Subsequent frames are sent to the region proposal network for classification and regression through the search branch to obtain the position with the largest response and its bounding box, in preparation for tracking subsequent frames, and finally completing the entire tracking task.

[0030] In order to verify the effectiveness of this embodiment, the UAV123 data set is selected as the test data for this experiment. The UAV123 data set is data collected by drones at low altitudes, with 123 video sequences, a total of more than 110,000 frames, and contains a variety of tracking scenes, such as people, ships, cars, and buildings. It involves changes in multiple attributes, such as 12 different types such as illumination changes, scale changes, fast movement, background blur, and occlusion. In the tracking process of drones, camera shake, variable scales, and inconsistent tracking scenes and camera shooting angles often occur, resulting in tracking difficulties and great challenges. The tracking performance is mainly evaluated by success rate and accuracy. The success rate refers to the ratio of the overlapping area of ​​the bounding box and the real marked bounding box greater than the set threshold to the total number of bounding boxes in the current video image, and the accuracy refers to the ratio of the center error of the bounding box from the real bounding box less than the set threshold to the total number of bounding boxes in the current video image.

[0031] In this experiment, when the intersection-and-union ratio is set to be greater than 0.6, it is considered a positive sample, and when it is less than 0.3, it is considered a negative sample. There are 1805 candidate bounding boxes calculated from the bounding box in a video image. Due to the large number, the total number of samples is limited to 256 in a set of training processes, and the ratio of the number of positive samples to the number of negative samples is 1 to 3. The test time of this embodiment in the drone data set is 1064 seconds, the frame rate is about 106 frames / second, the test time of the original SiamRPN is 1066 seconds, and the frame rate is about 106 frames / second. This embodiment is 2 seconds faster than the original SiamRPN. Compared with the original single target tracking algorithm SiamRPN, this embodiment improves the accuracy rate by 8.29%, the success rate by 11.63%, the accuracy rate by 5.87%, and the success rate by 10.9% under the conditions of background blur and illumination changes. It is also better than other mainstream target tracking algorithms because the strip pool module is added to the template branch, which can strengthen the dependence of spatial semantics on the backbone network and adapt to changes in illumination. At the same time, the context network block is added to the search branch to input the regression points of the region proposal network, which strengthens the network's understanding of the global context and makes the regressed bounding box more accurate. Therefore, this embodiment can achieve good tracking effects when the background is blurred and the lighting changes greatly.

[0032] The above description is only a preferred implementation manner of the present invention, but the protection scope of the present invention is not limited thereto, and any modification and replacement based on the technical solution and inventive concept provided by the present invention should be included in the protection scope of the present invention.

Claims

1. A twin region proposal network model device for unmanned aerial vehicle target tracking, characterized in that: It includes a template branch unit and a search branch unit, wherein the template branch unit includes a first convolution module, a strip pooling module, a second convolution module, a third convolution module, a first matching module and a first output module, and the search branch unit includes a fourth convolution module, a global context network module, a fifth convolution module, a sixth convolution module, a second matching module and a second output module; The first convolution module and the fourth convolution module constitute a twin network, the first convolution module is connected to the strip pooling module, and the fourth convolution module is connected to the global context network module; The second convolution module, the third convolution module, the first matching module, the first output module, the fifth convolution module, the sixth convolution module, the second matching module and the second output module constitute a region proposal network, the second convolution module and the third convolution module are both connected to the first matching module, and the first matching module is connected to the first output module; the fifth convolution module and the sixth convolution module are both connected to the second matching module, and the second matching module is connected to the second output module; wherein the strip pooling module is respectively connected to the second convolution module and the fifth convolution module, the fourth convolution module is connected to the third convolution module, and the global context network module is connected to the sixth convolution module; The strip pooling module is used to perform strip pooling operations. Strip pooling is performed along the horizontal and vertical directions of the window. The long-to-narrow kernel is used to establish a long-range context relationship. The expansion of the receptive field of the backbone network is beneficial to the classification of the target and background during the tracking process. The global context network module is used to establish the dependency relationship of the network's remote context and deepen the network's global understanding ability, automatically increase the proportion of channels related to the target features, and reduce the proportion of channels unrelated to the target features, change the dependency between different channels, and make the bounding box regression more accurate.

2. The twin region proposal network model device for unmanned aerial vehicle target tracking according to claim 1, characterized in that: The template branch unit input image size is 127×127×3, and the search branch unit input image size is 255×255×3.

3. The twin region proposal network model device for unmanned aerial vehicle target tracking according to claim 1, characterized in that: The network model uses distance intersection and union ratio to calculate the bounding box.

4. The twin region proposal network model device for unmanned aerial vehicle target tracking according to claim 3 is characterized in that: The classification loss function in the region proposal network adopts the cross entropy loss function, and the regression loss function is the L1 norm loss function.

Citation Information

Patent Citations

  • A target tracking method for guiding a twin network based on an intersection-over-union (IOU)

    CN112509008A

  • A computer vision application-oriented anchor-frame-free target tracking algorithm

    CN113554679A