UAV Perspective Target Tracking Method Based on Sequence Perception and Feature Enhancement
By building a tracking network with sequence perception and feature enhancement, video sequence features are extracted and image features are enhanced, and tracking failure caused by insufficient features of small and medium-sized target tracking of drone perspective targets and camera motion is solved, achieving higher tracking accuracy and stability.
Patent Information
- Application Number
- CN202211194947.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-27
AI Technical Summary
In the drone perspective target tracking, the network lacks the characteristics of small targets, and camera movement causes drastic changes in the position of the target in the image, resulting in tracking failure.
A tracking network of sequence perception and feature enhancement is constructed, video sequence features are extracted through sequence feature perception subnetwork, camera motion mode is learned, and image features are enhanced through image feature enhancement subnetwork to supplement the features of small targets.
It effectively solves the problems of tracking failure caused by insufficient characteristics of small targets and camera movements of drone perspective target tracking, and improves the tracking accuracy and stability of small targets.
Smart Images

Figure CN115661686B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and further relates to a method for tracking an object from an unmanned aerial vehicle (UAV) perspective based on sequence perception and feature enhancement in the technical field of object tracking. The present invention can use a UAV to track pedestrians, vehicles, etc. moving in urban security, autonomous driving, and intelligent transportation. Background Art
[0002] The task of object tracking from a UAV perspective is to predict the size and position of an object in subsequent frames given the size and position of the object in the initial frame of a given UAV perspective video sequence. There are two reasons for the failure of the object tracking task from a UAV perspective. First, since the UAV flies at a relatively high altitude during shooting, the size of the object being photographed is small in the image, the features are blurred and easily occluded by the environment. Second, in order to ensure that the object is continuously captured by the camera when the UAV shoots the object, the camera angle needs to be frequently adjusted, resulting in a situation where the object moves violently in the video.
[0003] Guizhou University disclosed in its patent document "A New UAV Object Tracking Algorithm" (Patent Application No.: CN114972439A, Publication No.: CN 110517291 A) a method of using a feature correlation network based on a self-attention mechanism for object tracking from a UAV perspective. This method first inputs a template image, a search region image, and a dynamic template update image into a feature extraction network, then inputs the output features into a network based on an attention mechanism to obtain global feature information, and finally inputs the template features and search region features into a feature correlation network for decoding to obtain the tracking result. The disadvantage of this method is that since this method does not consider the context information of the video sequence, it cannot solve the above-mentioned small object problem and camera movement problem. When the object size is small and the features are insufficient, and when the movement of the UAV camera causes a drastic change in the position of the object in the image, the tracker is likely to lose the object, resulting in tracking failure.
[0004] Goutam, Martin and others disclosed a sequence-based method for object tracking from an unmanned aerial vehicle (UAV) perspective in their published paper "Know Your Surroundings: Exploiting Scene Information for Object Tracking" (Conference Name: European Conference on Computer Vision’2020). This method first inputs the current frame, previous frame and previous state vector into a network model, then inputs the current frame into an image feature extraction network to obtain appearance features. At the same time, the current frame, previous frame and previous state vector are input into an aggregation network model together to update the previous state vector. Finally, the updated state vector and appearance features are input into a decoder to obtain the tracking result. Since the previous state vector used in this method contains sequence information, to a certain extent, this method solves the deficiency that the method in the patent document "A New UAV Object Tracking Algorithm" applied by Guizhou University does not utilize the context information of the video sequence. However, the deficiency that still exists in this method is that it cannot obtain accurate tracking results when dealing with small object problems and camera motion problems in the process of UAV perspective object tracking, indicating that the design of the tracking network structure of this method limits its extraction and characterization of sequence features. Summary of the Invention
[0005] The object of the present invention is to address the deficiencies of the above-mentioned existing technologies, and propose a UAV perspective object tracking method based on sequence perception and feature enhancement, which is used to solve the problem of insufficient features of small objects by the network during UAV perspective object tracking, and the problem of tracking failure caused by the drastic change in the position of the object in the image due to camera movement.
[0006] The idea for achieving the object of the present invention is as follows: construct a tracking network for sequence perception and feature enhancement. First, use a sequence feature perception sub-network to extract video sequence features, and learn the motion pattern of the camera according to the implicit camera motion information contained in the sequence features. Then use an image feature enhancement sub-network to enhance the image features, and use the sequence features to supplement the features of small objects in the time series dimension. Thus, the camera motion problem and small object problem in the process of UAV perspective object tracking are solved.
[0007] The implementation steps of the present invention are as follows:
[0008] Step 1, construct a tracking network for sequence perception and feature enhancement:
[0009] Step 1.1, construct a sequence feature calculation sub-network;
[0010] Build a sequence feature calculation sub-network composed of a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. Among them, the first convolutional layer and the second convolutional layer are connected in series to form a first series module, the third convolutional layer and the fourth convolutional layer are connected in series to form a second series module, and then the first series module and the second series module are connected in parallel to form a sequence feature calculation sub-network. Set the number of convolutional kernels of the first to fourth convolutional layers to 64, and the convolutional kernel size to 1×1;
[0011] Step 1.2, construct an image feature enhancement sub-network;
[0012] Build an image feature enhancement sub-network composed of a first convolutional layer, a second convolutional layer, and a third convolutional layer. Among them, the first convolutional layer, the second convolutional layer, and the third convolutional layer are connected in series in sequence to form an image feature enhancement sub-network. Set the number of convolutional kernels of the first to third convolutions to 1, 1, and 512 in sequence, and the convolutional kernel size to 3×3, 3×3, and 1×1 in sequence;
[0013] Step 1.3, connect the image feature extraction sub-network Resnet50 and the sequence feature calculation sub-network in parallel, and then connect them in series with the image feature enhancement sub-network and the DiMP tracking model decoder in sequence to form a tracking network for sequence perception and feature enhancement;
[0014] Step 2, generate a training set;
[0015] Step 2.1, randomly collect at least 100,000 image pairs from natural image videos. Each pair of image pairs consists of two adjacent frames in the video, and there is the same target of any type in both frames. The two frames are respectively called the current frame image and the previous frame image;
[0016] Step 2.2, preprocess each pair of image pairs. Take the center of the target in the current frame image of each pair of image pairs as the cropping center, crop the sizes of the two images in the image pair to [288×288×3], and then perform random translation or scaling processing on the cropped image pair;
[0017] Step 2.3, for each pair of preprocessed image pairs, label the position of the target in the current frame image in the form of a position box. The position box is represented by [c, w, h], where c, w, and h are the center, width, and height of the position box. Take [c, w, h] as the ground truth position box of the target of the image pair, and at the same time generate a Gaussian label map with c as the center as the target classification label map of the image pair;
[0018] Step 2.4, form a training set with all preprocessed image pairs and their corresponding target classification label maps and ground truth position boxes;
[0019] Step 3, train the network:
[0020] Input the image pairs in the training set into the tracking network with sequence perception and feature enhancement. Use the SGD optimization algorithm to iteratively train the parameters of the tracking network with sequence perception and feature enhancement offline until the loss function converges, and obtain the trained network.
[0021] Step 4, use the trained network to perform online tracking on the target in the UAV perspective video.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] First, through the constructed sequence feature perception sub-network, the present invention solves the shortcoming of insufficient representation of sequence information in the existing methods, enabling the present invention to learn camera motion information based on sequence features and having the advantage of not losing the tracking target even after the UAV camera moves.
[0024] Second, through the constructed image feature enhancement sub-network, the present invention solves the shortcoming of insufficient representation of small targets in the existing methods, enabling the present invention to supplement the appearance features of small targets using sequence features and having the advantage of not losing the tracking target due to insufficient appearance features when tracking small targets. Description of the Drawings
[0025] Figure 1 is the flowchart of the present invention;
[0026] Figure 2 is the simulation result diagram of the present invention. Detailed Embodiments
[0027] The following will further describe the present invention in detail with reference to the drawings and embodiments.
[0028] Refer to Figure 1 and the embodiments to further describe the specific implementation steps of the present invention in detail.
[0029] Step 1, construct a tracking network with sequence perception and feature enhancement.
[0030] Step 1.1, construct a sequence feature calculation sub-network.
[0031] Build a sequence feature calculation sub-network composed of a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. Among them, the first convolutional layer and the second convolutional layer are connected in series to form a first series module, the third convolutional layer and the fourth convolutional layer are connected in series to form a second series module, and then the first series module and the second series module are connected in parallel to form a sequence feature calculation sub-network. Set the number of convolutional kernels of the first to fourth convolutional layers to 64, and the convolutional kernel size to 1×1.
[0032] Step 1.2, construct an image feature enhancement sub-network.
[0033] Build an image feature enhancement sub-network consisting of a first convolutional layer, a second convolutional layer, and a third convolutional layer. Among them, the first convolutional layer, the second convolutional layer, and the third convolutional layer are connected in series in sequence to form the image feature enhancement sub-network. Set the number of convolutional kernels of the first to third convolutions to 1, 1, and 512 in sequence, and the convolutional kernel sizes to 3×3, 3×3, and 1×1 in sequence.
[0034] Step 1.3, connect the image feature extraction sub-network Resnet50 and the sequence feature calculation sub-network in parallel, and then connect them in series with the image feature enhancement sub-network and the DiMP tracking model decoder in sequence to form a sequence-aware and feature-enhanced tracking network.
[0035] The above-mentioned image feature extraction sub-network Resnet50 is composed of a first convolutional layer, a first residual network module, a second residual network module, and a third residual network module connected in series in sequence. The DiMP tracking model decoder is composed of a classification prediction module and a position regression module connected in parallel.
[0036] Step 2, generate a training set.
[0037] Step 2.1, randomly collect 100,000 image pairs from the natural image video datasets GOT10K, LaSOT, ImageNet, and TrackingNet. Each pair of image pairs consists of two adjacent frames in the video, and there is the same target of any category in both frames. The two frames are respectively called the current frame image and the previous frame image.
[0038] Step 2.2, preprocess each pair of image pairs. Take the center of the target in the current frame image of each pair of image pairs as the cropping center, crop the sizes of the two images in the image pair to [288×288×3], and then perform random translation or scaling processing on the cropped image pair.
[0039] Step 2.3, for each pair of preprocessed image pairs, label the position of the target in the current frame image in the form of a position box. The position box is represented by [c, w, h], where c, w, and h are the center, width, and height of the position box. Take [c, w, h] as the ground truth position box of the image pair, and at the same time generate a Gaussian label map with c as the center as the target classification label map of the image pair.
[0040] Step 2.4, form a training set with all preprocessed image pairs and their corresponding target classification label maps and ground truth position boxes.
[0041] Step 3, train the network.
[0042] Step 3.1: Input the image pairs in the training set into the tracking network with sequence perception and feature enhancement. Use the sequence feature calculation sub-network to calculate the sequence features of the current frame image and the previous frame image, then use the image feature enhancement sub-network to enhance the image features of the current frame image using the sequence features. Finally, use the DiMP tracking model decoder to decode the enhanced features to obtain the predicted classification response map and the predicted target position box.
[0043] Take the image pair I c and I p as an example of the propagation process in the tracking network with sequence perception and feature enhancement to further illustrate Step 3.1.
[0044] Input I c and I p into the first convolutional layer of the image feature extraction sub-network Resnet50 respectively. This convolutional layer outputs the first convolutional product features of the current frame image and the first convolutional product features of the previous frame image where represent the size of the feature map, and width, height, and channel represent the width, height, and number of channels of the feature map respectively. Then input into the first residual network module, the second residual network module, and the third residual network module in sequence to obtain the second residual network module features of the current frame image and the second residual network module features of the current frame image
[0045] According to and calculate the sequence features. Input into the first convolutional layer and the second convolutional layer of the sequence feature calculation sub-network in sequence to obtain the current frame key feature At the same time, input into the third convolutional layer and the fourth convolutional layer of the sequence feature calculation sub-network in sequence to obtain the previous frame key feature Transpose and multiply it with to obtain the initial sequence feature Sum the element values in in the channel dimension and take the average to obtain the final sequence feature
[0046] Enhance the image features according to the sequence feature F c,p Input F c,p into the first convolutional layer, the second convolutional layer, and the third convolutional layer of the image feature enhancement sub-network respectively. The second convolutional layer and the third convolutional layer output the first downsampling feature and the second downsampling feature Then input and Separate from and features are concatenated to obtain the initialized enhanced feature and Then and are respectively input into the fourth convolutional layer module and the fifth convolutional layer module of the image feature enhancement sub-network to obtain the enhanced features and
[0047] The enhanced feature is input into the decoder classification prediction module of the DiMP tracking model to obtain the predicted classification response map, and then and are simultaneously input into the decoder position regression module of the DiMP tracking model to obtain the predicted target position box.
[0048] Step 3.2, using the SGD optimization algorithm, offline iteratively train the parameters of the tracking network for sequence perception and feature enhancement until the loss function converges, and obtain the trained network.
[0049] The loss function L is as follows:
[0050]
[0051] Among them, n represents the total number of image pairs in the training set, i represents the serial number of the image pair in the training set, ∩ represents the intersection operation, ∪ represents the union operation, y cls and respectively represent the target classification label map and the classification response map predicted by the network, y bbox and represent the target true position box and the target position box predicted by the network.
[0052] Step 4, use the trained network to perform online tracking on the target in the UAV perspective video.
[0053] Step 4.1, determine the target to be tracked according to the given target position and target size in the first frame of the UAV perspective video.
[0054] Step 4.2, starting from the second frame, with the center of the target position in the previous frame as the cropping center, crop the current frame and the previous frame to [288×288×3] and input them into the trained tracking network for sequence perception and feature enhancement to obtain the predicted classification response map and the predicted target position box; if all the values in the predicted classification response map are greater than 0.6, it is regarded as successful target tracking, and step 4.3 is executed; if all the values in the predicted classification response map are less than 0.6, otherwise it is regarded as failed target tracking, and 4.4 is executed.
[0055] Step 4.3, form a training pair by combining the enhanced features, the predicted classification response map, and the predicted target location box of the current frame image, and store the training pair in the memory feature set Set suc 。
[0056] Step 4.4, according to the training pairs in the memory feature set Set suc sort the training pairs in descending order according to the values of the predicted classification response map in the training pairs. Use the first 20 sorted training pairs to train the parameters of the DiMP tracking model decoder. If the number of training pairs in the memory feature set Set suc is less than 20, supplement with the first sorted training pair.
[0057] Step 4.5, determine whether the current frame is the last frame of the tracking video. If so, execute Step 4.2; otherwise, complete the target tracking process.
[0058] The following further describes the effect of the present invention in combination with simulation experiments:
[0059] 1. Simulation experiment conditions:
[0060] The software platform for the simulation experiment of the present invention is: The simulation of the present invention is carried out on the Ubuntu18.04 system with an Intel(R) Core(TM) i7-11700k CPU, an Nvidia 3090 graphics card, and 32G of memory, using the Python3.7 software and the Pytorch1.8 deep learning toolkit.
[0061] The simulation experiment of the present invention uses 50 video sequences in the self-collected UAV perspective dataset to test the target tracking method. In the above dataset, 68% of the video sequences have small-sized targets (the number of target pixel values is less than 900), and 52% of the video sequences have camera movement.
[0062] 2. Simulation content and its result analysis:
[0063] The simulation experiment of the present invention uses the present invention and two existing technologies (the Kys method of tracking method based on sequence information and the DiMP method of tracking method based on discriminant prediction model) to track the targets in the above 50 video sequences respectively, and the results are as Figure 2 shown. Figure 2 Only show the simulation results of 6 key frames in a representative video sequence among the 50 video sequences. The 6 key frames are the 1st frame, the 40th frame, the 56th frame, the 93rd frame, the 196th frame, and the 203rd frame respectively. The frame numbers are marked in the upper left corner of each frame image.
[0064] In the simulation experiment of the present invention, the two existing technologies adopted refer to:
[0065] The Kys method, a tracking method based on sequence information in the prior art, refers to a method based on sequence information that can be used for target tracking from a drone's perspective, which was disclosed in the paper "Know Your Surroundings: Exploiting Scene Information for Object Tracking" (Conference Name: European Conference on Computer Vision'2020) by Goutam, Martin et al.
[0066] The prior art tracking method based on the discriminative prediction model DiMP method refers to a method based on sequence information that can be used for drone perspective target tracking disclosed in the paper "Learning Discriminative Model Prediction for Tracking" (Conference Name: Computer Vision and Pattern Recognition'2019) published by Goutam and Martin et al.
[0067] Combine the following Figure 2 The simulation diagram of the present invention is further described.
[0068] Figure 2 The green position frame in the figure represents the position frame of the real target before simulation, and the red, yellow and blue position frames represent the target position frames predicted by the method of the present invention, the Kys method and the DiMP method respectively. Figure 2 It can be seen that although the method of the present invention, the Kys method and the DiMP method can track the target in the 1st, 40th, 56th, 93rd and 196th frames of the video sequence, the camera moves in the 196th frame image, and the above three methods temporarily lose the target. It can be seen from the 203rd frame image that due to the violent camera movement and the small size of the target, the Kys method and the DiMP method cannot track the target in subsequent frames. Only the method of the present invention tracks the target after the camera is stabilized.
[0069] In order to verify the simulation results of the present invention, two evaluation indicators (prediction center distance CLE and prediction position box overlap area IoU) are used to evaluate the target tracking results of the above three methods respectively. Using the following formula, CLE and IoU calculate the center distance and overlap area of the position box of the real target and the target position box predicted by the three methods, and all the calculation results are plotted in Table 1:
[0070]
[0071]
[0072] Among them, frame is the total number of frames of 50 video sequences in the self-built UAV perspective dataset, and are the network prediction values of the abscissa and ordinate of the center of the target position in the video frame with serial number f, respectively, and are the true values of the abscissa and ordinate of the center of the target position in the video frame with serial number f, respectively. and are the network prediction value and the true value of the position box of the target in the video frame with serial number f.
[0073] Table 1. Quantitative analysis table of the target tracking results of the present invention and advanced technologies in the simulation experiment
[0074] CLE IoU KYS 91.36 0.65 DiMP 28.76 0.68 The method of the present invention 20.55 0.71
[0075] Combined with Table 1, it can be seen that the target tracking results of the present invention are improved compared with the Kys method and the DiMP method, which proves the advancement of the method of the present invention.
[0076] The above simulation experiments show that: the UAV perspective target tracking method based on sequence perception and feature enhancement of the present invention constructs a sequence feature calculation sub-network and an image feature enhancement sub-network, extracts video sequence features to learn the motion mode of the UAV camera, and uses sequence features to supplement the features of small targets in the time series dimension. Thus, the camera motion problem and small target problem in the process of UAV perspective target tracking are solved.
Claims
1. A method for tracking an object from an unmanned aerial vehicle (UAV) perspective based on sequence perception and feature enhancement, characterized in that, Construct a sequence feature calculation sub-network and an image feature enhancement sub-network respectively; the specific steps of this tracking method are as follows: Step 1, construct a tracking network for sequence perception and feature enhancement; Step 1.1, construct a sequence feature calculation sub-network; Build a sequence feature calculation sub-network composed of a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer. Among them, the first convolutional layer and the second convolutional layer are connected in series to form a first series module, the third convolutional layer and the fourth convolutional layer are connected in series to form a second series module, and then the first series module and the second series module are connected in parallel to form a sequence feature calculation sub-network. Set the number of convolutional kernels of the first to fourth convolutional layers to 64, and the convolutional kernel size to 1×1; Step 1.2, construct an image feature enhancement sub-network; Build an image feature enhancement sub-network composed of a first convolutional layer, a second convolutional layer, and a third convolutional layer. Among them, the first convolutional layer, the second convolutional layer, and the third convolutional layer are connected in series in sequence to form an image feature enhancement sub-network. Set the number of convolutional kernels of the first to third convolutions to 1, 1, and 512 in sequence, and the convolutional kernel size to 3×3, 3×3, and 1×1 in sequence; Step 1.3, connect the image feature extraction sub-network Resnet50 and the sequence feature calculation sub-network in parallel, and then connect them in series with the image feature enhancement sub-network and the DiMP tracking model decoder in sequence to form a tracking network for sequence perception and feature enhancement; Step 2, generate a training set; Step 2.1, randomly collect at least 100,000 image pairs from natural image videos. Each pair of image pairs consists of two adjacent frames in the video, and there is the same target of any type in both frames. The two frames are respectively called the current frame image and the previous frame image; Step 2.2, preprocess each pair of image pairs. Take the center of the target in the current frame image of each pair of image pairs as the cropping center, crop the sizes of the two images in the image pair to [288×288×3], and then perform random translation or scaling processing on the cropped image pair; Step 2.3, for each pair of preprocessed image pairs, label the position of the target in the current frame image in the form of a position box. The position box is represented by [c, w, h], where c, w, and h are the center, width, and height of the position box. Take [c, w, h] as the ground truth position box of the image pair, and at the same time generate a Gaussian label map with c as the center as the target classification label map of the image pair; Step 2.4, form a training set with all preprocessed image pairs and their corresponding target classification label maps and ground truth position boxes; Step 3, train the network; Input the image pairs in the training set into the tracking network for sequence perception and feature enhancement, use the SGD optimization algorithm, and iteratively train the parameters of the tracking network for sequence perception and feature enhancement offline until the loss function converges to obtain a trained network; Step 4, use the trained network to perform online tracking on the target in the UAV perspective video.
2. The method for tracking an object from an unmanned aerial vehicle perspective based on sequence perception and feature enhancement according to claim 1, wherein: The loss function L described in Step 3 is as follows: Among them, n represents the total number of image pairs in the training set, i represents the serial number of the image pair in the training set, ∩ represents the intersection operation, ∪ represents the union operation, y cls and respectively represent the target classification label map and the classification response map predicted by the network, y bbox and represent the target true position box and the target position box predicted by the network.
Citation Information
Patent Citations
Road vehicle tracking method based on multi-feature space fusion
CN110517291A
Novel unmanned aerial vehicle target tracking algorithm
CN114972439A