A single-target vehicle tracking method based on UAV aerial photography
By adopting multi-feature fusion image matching and K++ neighborhood search algorithm in drone aerial videos, combined with the anti-occlusion strategy of vehicle motion state estimation, the accuracy and speed problems of single-target tracking in drone aerial videos are solved, and fast and accurate single-target vehicle tracking is achieved.
Patent Information
- Application Number
- CN202210156746.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The existing single-target tracking algorithm based on twin networks has problems such as insufficient sample results in drone aerial videos, which are difficult to meet real-time tracking speed.
The image matching algorithm based on multi-feature fusion is adopted, combined with K++ neighborhood search and anti-occlusion algorithm based on vehicle motion state estimation, the image matching accuracy is improved through the fusion of color histogram features and HOG features, and the target position is predicted using the vehicle motion state when the target is blocked.
It realizes fast and accurate single-target vehicle tracking in drone aerial videos, with good versatility and scalability, and improves tracking accuracy and speed.
Smart Images

Figure CN114529584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a single-target vehicle tracking method based on UAV aerial photography. Background Art
[0002] Target tracking tasks can be divided into two categories: one is multi-target tracking, and the other is single-target tracking. Multi-target tracking refers to the task of tracking all targets or all individuals of a certain type of target in a video sequence. It not only involves continuous tracking of each target, but also involves the recognition of different targets, self-occlusion and mutual occlusion processing, and the correlation of detection results with tracking results. Compared with multi-target tracking, we are more inclined to single-target tracking because in a video sequence, we often focus more on the movement of a certain individual that appears. Therefore, the development of single-target tracking technology is very popular and extremely rapid. From the traditional methods of processing between video frames, representative ones such as static background, frame difference method, optical flow method, Meanshift and Camshift algorithms belong to early single-target tracking methods, which have the characteristics of many practical applications, high FPS, and low requirements for the computing power of devices. Then, the combination of detection and tracking began to appear. Machine learning-based methods extract image features and train classifiers for classification (such as SVM), so that the trained classifier can find the optimal region in the next frame. One method is based on the generative model, and the other is the discriminative model, collectively referred to as detection and tracking. The current most representative single-target tracking method is the kernel correlation filtering method and the algorithm combined with deep learning on this basis. It has refreshed the accuracy and speed of single-target tracking again. However, because the deep algorithm has a large amount of computation and requires high computing power of the device, many tracking algorithms use online fine-tuning, so the speed is not very ideal, and the practical application has certain limitations, and there is still a large room for development.
[0003] In recent years, single-object tracking algorithms based on deep learning have been developed. The SiamFC algorithm based on Siamese networks proposed in CVPR2016 has pushed single-object tracking into a new era. By leveraging the characteristics of Siamese networks, after extracting features of the template and the image through the same network, the feature vector of the template is cross-correlated with the feature vector of the search image to obtain a response map, and the position with the maximum response is the position of the target. Almost all subsequent single-object tracking algorithms based on deep learning are developed based on this algorithm. For example, the SiamMask algorithm adds a simple 1×1 convolutional kernel with 2 channels on the basis of cross-correlation to obtain outputs of two branches for different task processing; another representative SiamRPN algorithm, based on SiamFC, integrates the classification and regression of the RPN network in Faster-RCNN with SiamFC, not only improving the tracking accuracy but also significantly enhancing the tracking speed. The latest Dimp (Learning discriminative model prediction for tracking) tracking algorithm released by the Martin Daniel Laboratory is based on SiamFC, adds an online training classifier, and optimizes the classifier using the information of the previous and subsequent frames of the video to adjust the model in real time, improving the tracking accuracy.
[0004] Although the current single-object tracking algorithms based on Siamese networks have relatively good tracking effects, since all the information obtained by the network is provided by the first frame, the amount of information obtained is too small. Therefore, aiming at the problems of low accuracy caused by insufficient samples and difficulty in meeting real-time requirements in the current object tracking field, Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a single-object vehicle tracking method based on UAV aerial photography, which can quickly and accurately perform single-object tracking on a certain vehicle target in the video captured by the UAV, and has good versatility and scalability.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is:
[0007] A single-object vehicle tracking method based on UAV aerial photography, comprising the following steps:
[0008] Step 1: Load the UAV aerial photography video that needs vehicle tracking, pause at the first frame, manually select the target vehicle to be tracked with the mouse, and the selected area is the area to be tracked. The target vehicle within the area to be tracked is the tracking target; then starting from the second frame of the video, use the object detection algorithm to detect the target vehicle that appears in the video frame;
[0009] Step 2: Determine whether the target vehicle in the current frame is completely blocked. If not, proceed to step 3; otherwise, proceed to step 5;
[0010] Step 3: Establish a K++ neighborhood around the tracking box, use the K++ neighborhood to filter redundant target detection boxes, keep the vehicles that may be the tracking target in the K++ neighborhood, and then calculate the IoU of the tracking box and the detection box and the offset of the center point;
[0011] Step 4: Extract the target in the filtered detection frame as an image, match it with the tracking target selected in the first frame using a multi-feature fusion image matching algorithm, calculate the image similarity between the tracking target and the filtered target and sort them, and then combine the calculation results of step 3 to comprehensively determine which detected target in the current frame is the target to be tracked, and then update the tracking frame; execute step 6;
[0012] Step 5: When the tracking target is occluded, an anti-occlusion algorithm based on vehicle motion state estimation is used. After the tracking starts for more than 20 frames, the average speed of the vehicle in the video is recorded every 20 frames. According to this method, when the target disappears from the field of view, the coordinates at the time of disappearance are saved, and the target's moving speed is stopped and saved in the previous 20 frames. If the target disappears within 50 frames, the moving trajectory and coordinates within the disappearance are estimated normally, and the K neighborhood of the current estimated position is obtained at the same time, and the position where the target may appear after 50 frames is recorded. The K neighborhood is set at this position to wait for the target to be captured. If the vehicle is re-detected and captured by the K neighborhood, the image matching algorithm in step 4 is called to start matching. If the match is successful, the tracking continues. If the match is not successful, the full image matching is turned on, the movement speed and coordinates are released, and the tracker searches by itself. The tracking method of steps 3-4 is called, and the multi-feature fusion matching algorithm is used to match the target with the greatest similarity to the initially selected tracking target that appears in the current field of view, and the K++ neighborhood is re-established based on the coordinates of the target; execute step 6;
[0013] Step 6: Determine whether the video has ended; if so, end the detection; otherwise, receive the next frame and return to step 2.
[0014] The beneficial effects of adopting the above technical solutions are as follows: The single-target vehicle tracking method based on UAV aerial photography provided by the present invention realizes single-target tracking by using an image matching algorithm based on multi-feature fusion, a target prediction K++ neighborhood search algorithm, and an anti-occlusion algorithm based on vehicle motion state estimation. The image matching algorithm based on multi-feature fusion combines the image color histogram feature and the HOG feature. By this method, the representation degree of the image using only a single feature and the accuracy rate in the image matching process can be significantly improved. The target prediction K++ neighborhood search algorithm screens the results of target detection, which helps to reduce the calculation amount, and compared with the original K neighborhood search algorithm, this algorithm has higher accuracy in searching for the tracking target within the neighborhood and can more effectively eliminate the interference caused by the appearance of similar targets. When the tracking target is completely occluded during the tracking process, the anti-occlusion algorithm based on vehicle motion state estimation is adopted. When the motion state of the vehicle changes little during the occlusion period compared with that before occlusion, the algorithm calculates the moving speed and coordinates of the target before being occluded in the video, predicts the motion state of the target during the disappearance period, and thus estimates the possible position where the target may appear after driving out of the occluder, realizing the repositioning of the target. The method of the present invention can quickly and accurately perform single-target tracking on a certain vehicle target in the video captured by the UAV, and has good versatility and scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flowchart of the single-target vehicle tracking method based on UAV aerial photography provided by an embodiment of the present invention;
[0016] Figure 2 It is a method for selecting a tracking target provided by an embodiment of the present invention; among them, Fig. (2a) is an incorrect box selection method, and Fig. (2b) is a correct box selection method;
[0017] Figure 3 It is the tracking effect when a small part of occlusion occurs provided by an embodiment of the present invention; among them, Fig. (3a), Fig. (3b), and Fig. (3c) are the images of the 370th frame, the 400th frame, and the 420th frame respectively;
[0018] Figure 4 It is the situation where similar vehicles appear provided by an embodiment of the present invention; among them, Fig. (4a), Fig. (4b), and Fig. (4c) are the images of the 146th frame, the 219th frame, and the 241st frame respectively;
[0019] Figure 5 It is the situation where the target is completely occluded provided by an embodiment of the present invention; among them, Fig. (5a), Fig. (5b), and Fig. (5c) are the images of the 134th frame, the 140th frame, and the 144th frame respectively. DETAILED DESCRIPTION OF THE INVENTION
[0020] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0021] Target tracking is an important task in computer vision. From traditional tracking algorithms to contemporary target tracking based on deep learning, countless scholars and researchers have invested a lot of time and energy to make very important contributions to the target tracking task. Today's computer vision neighborhood, especially the target tracking direction based on deep learning, is the most popular. The present invention also starts from the deep learning target detection network YOLOv4, combining image matching and target prediction to achieve single target tracking. For single target tracking, the point of view proposed by the present invention is that we must first detect the target and know the type of the target, and then add some means between each video frame to link the same target between frames and find and mark it in the target category to which it belongs, so as to achieve single target tracking. Therefore, the main contents of the present invention are an image matching algorithm based on multi-feature fusion, a target prediction K++ neighborhood search algorithm and an anti-occlusion algorithm based on vehicle motion state estimation, which are divided into the following three points:
[0022] (1) Multi-feature fusion image matching algorithm IMF. Starting with the basic features of the image, the method of calculating the similarity of each basic feature is tested to find the best way to match the target when the drone is shooting a vehicle. After finding that the color histogram feature is more suitable for vehicle tracking, the similarity calculation method of this feature is stress tested by selecting a road section with heavy traffic. It is found that a single image feature cannot represent the target well, so the vehicle contour feature HOG feature is fused, and the corresponding weight is added to the obtained similarity to obtain the final similarity score. After comprehensive judgment, the target is screened. Therefore, taking the basic features of the image as the basis, an image matching algorithm with multi-feature fusion is proposed.
[0023] (2) Target prediction K++ neighborhood search algorithm. First, the classification algorithm KNN is studied, and then the K neighborhood search algorithm proposed based on its classification idea is introduced. On this basis, the K++ neighborhood search algorithm is proposed. This algorithm integrates the K neighborhood search algorithm with the idea of IoU and the Euclidean distance between the center point of the tracking box and the prediction box to obtain the K++ neighborhood search algorithm. Experiments show that this algorithm can better cope with the situation when the same vehicle as the tracking target appears than the original K neighborhood search algorithm. This algorithm can still ensure the correct tracking of the target.
[0024] (2) An anti-occlusion algorithm AOE based on vehicle motion state estimation. This algorithm starts recording the vehicle's motion information, including average speed and displacement direction, every 20 frames after the target motion exceeds 20 frames. When the target is completely occluded, it saves the motion information of the first 20 frames and starts estimating the target's motion trend during the occlusion. Experiments prove that the AOE anti-occlusion algorithm can handle the situation of the reappearance of full occlusion of the target to a certain extent.
[0025] As Figure 1 shown, the single-target vehicle tracking method based on UAV aerial photography in this embodiment is specifically described as follows.
[0026] Step 1: Obtain the UAV aerial photography video. After the video is loaded, it pauses at the first frame, and the target to be tracked is selected by mouse in the first frame of the video. The selected area is the area to be tracked, and the target vehicle in the area to be tracked is the tracking target. Then, in the subsequent video frames, the tracking algorithm continuously finds the position of the target in the video. When selecting the target, try to select the target itself as much as possible. As shown in Figure (2b), do not include too much information outside the target, so that the apparent features of the target can be extracted more accurately. As shown in Figure (2a), too much information outside the target is selected. This part is implemented by OpenCV, and the callback function is used by listening to mouse events.
[0027] Step 2: Determine whether the target vehicle in the current frame is completely occluded. If not, continue to Step 3; if so, execute Step 5.
[0028] Step 3: Apply the K++ neighborhood search algorithm to predict the position of the tracking target.
[0029] After selecting the tracking target, it is necessary to use the target detection algorithm to detect all vehicles appearing in the second and subsequent frames of the video, and the K++ neighborhood search algorithm is needed to process the detection results. In order to enhance the prediction effect of the target prediction algorithm, more rigorous limiting conditions need to be added on the basis of the existing K neighborhood to improve its performance. In this embodiment, the K++ neighborhood search algorithm is improved by combining the ideas of IoU and center point offset, and a K++ neighborhood search algorithm is designed to predict the position of the tracking target.
[0030] The execution process of the K++ neighborhood search algorithm is as follows:
[0031] Step 3.1: According to the size of the tracking box in the previous frame, calculate the K neighborhood range corresponding to the tracking box when k = 2, and narrow the detection range of the current frame to this K neighborhood; the K neighborhood should satisfy Equation (1):
[0032]
[0033] where, W k and H kis the width and height of the K-neighborhood search area, W and H are the width and height of the target tracking box in the previous frame, and k is the aspect ratio of the two;
[0034] Step 3.2: If only one target is detected within this K-neighborhood in the current frame, that is, at least two-thirds of the area of the target's detection box is within the K-neighborhood range, then this target is the target in the previous frame. Update the tracking box and continue to execute Step 4; if more than two targets appear within this K-neighborhood in the current frame, then execute Step 3.3;
[0035] Step 3.3: Calculate the similarity between the targets within the K-neighborhood and the tracking target respectively to obtain similarity scores and sort them;
[0036] Step 3.4: Calculate the IoU and the Euclidean distance of the center points between the target detection box corresponding to the sorted similarity scores and the tracking box in the previous frame; Calculate the IoU to satisfy Equation (2):
[0037]
[0038] Among them, gt is the tracking box in the previous frame; bb (bounding box) is the detection box that appears within the K-neighborhood range in the current frame. Calculate the IoU using gt and the detection boxes within the K-neighborhood respectively, and select the detection box with the largest IoU value for retention, satisfying Equation (3):
[0039] IoU(gt,bb) max =Max(IoU(gt,bb1),...,IoU(gt,bb n )) (3)
[0040] Among them, IoU is the intersection over union of the detection box and the tracking box in the previous frame, gt is the tracking box in the previous frame, bb is the detection box within the K++-neighborhood in the current frame, and n is the number of detection boxes within the K++-neighborhood;
[0041] Calculate the Euclidean distance of the center points to satisfy Equation (4):
[0042]
[0043] Among them, d is the Euclidean distance between two points, c is the center point of the box, and its coordinates are (x, y); c gt is the center point of the tracking box in the previous frame, c bb is the center point of the detection box in the current frame; Select the detection box corresponding to the center point with the smallest Euclidean distance, and combine it with the similarity calculated by the previous image matching and the largest IoU to determine which detection box detects the tracking target.
[0044] The judgment order at this time is as follows: first, compare the similarity of the images, then exclude similar vehicles according to IoU, and finally select the tracking target using the Euclidean distance of the center points.
[0045] Step 4: After the detection boxes are filtered in Step 3, it is necessary to match the tracking target with the targets in the filtered detection boxes to select the position where the tracking target appears in the current frame. The multi-feature fusion image matching algorithm is used for image matching, and its execution process is as follows:
[0046] Step 4.1: After selecting the target to be tracked, extract the color histogram features and HOG features of the tracking target, and convert the two features into feature vectors.
[0047] Step 4.2: In subsequent frames, extract the detected targets of the same category as pictures, and also extract the color histogram features and HOG features of each target to obtain feature vectors. Regarding the calculation method of the color histogram feature vector, considering that if each primary color can take 256 (0 - 255) values, there are 16 million colors (256 to the power of 3) in the entire color space, resulting in a very large amount of calculation. Therefore, the range 0 - 255 is divided into four equally sized regions: [0, 63] is region 0, [64, 127] is region 1, [128, 191] is region 2, and [192, 255] is region 3. Each primary color value has four possible values, so there are 64 color types (4 to the power of 3) composed of three primary colors. Therefore, any color that appears in the image will definitely belong to one of these four regions. Then, count the number of pixels in each region. After such partitioning, the calculation amount is reduced to the greatest extent while obtaining a 64-dimensional feature vector.
[0048] Step 4.3: Calculate the color histogram feature similarity and HOG feature similarity between the tracking target and all the targets obtained in Step 4.2 respectively, and perform weighted scoring. Finally, sort all the scores. The calculation method of the color histogram feature similarity is cosine similarity, which satisfies Equation (5):
[0049]
[0050] The smaller the angle between two vectors in space, the more similar the directions they point to, which means they are more similar, and the greater the similarity. This exactly corresponds to the characteristic of the cosine function graph that the smaller the angle θ, the larger the corresponding function value. Therefore, based on the above theoretical basis, the size of the vector angle in the coordinate system space can be used as the judgment basis for the similarity degree of vectors. The smaller the angle, the larger the cosine value corresponding to the angle, indicating that they are more similar. Correspondingly, the calculation method of cosine also holds for multi-dimensional vectors. Assume that P and Q are two multi-dimensional vectors, P is [P1, P2,..., P m , Q is [Q1, Q2,..., Qm , where m is the vector dimension.
[0051] The similarity of HOG features uses the HOG feature descriptor, i.e., the feature vector, and finally calculates the Euclidean distance between the feature vectors. The smaller the distance, the more similar the two images are. Since the HOG feature vector is an n-dimensional vector, the corresponding Euclidean distance satisfies Equation (6):
[0052]
[0053] where d is the Euclidean distance of the vector, and x i , y i are the two coordinate values of the vector in the multi-dimensional space. It can be seen from the formula that the closer the values at the corresponding indices of the two feature vectors are, the smaller the distance d is, which means the more similar the two images are.
[0054] An excitation coefficient is given to the similarity of the color histogram feature and the similarity of the HOG feature respectively, with the aim of obtaining the most suitable calculation method for calculating the similarity between vehicles. In this embodiment, according to the experimental data and the application scenario of the algorithm, the weight of the color histogram feature is set to 1, and the weight of the HOG feature is set to 2. The candidate box image with the largest similarity is selected as the target to be tracked.
[0055] After calculating the cosine similarity of the color histogram feature and the Euclidean distance similarity of the HOG feature between the features of each candidate box that has been screened and participated in the calculation and the tracking target, they are respectively multiplied by their corresponding weights, and added to obtain the final similarity score that satisfies Equation (7):
[0056] S i = W1(S(c i , c t )) + W2(S(h i , h t )) (7)
[0057] In the formula, S is the similarity calculation function of the parameters in the parentheses, and S i is the total similarity score between the image in the i-th candidate box and the tracking target; W1 is the similarity weight coefficient of the color histogram feature, with a value of 1; W2 is the similarity weight coefficient of the HOG feature, with a value of 2; S(c i , c t ) is the similarity of the color histogram feature between the i-th candidate box and the tracking target t, S(h i , h t ) is the similarity of the HOG feature between the i-th candidate box and the tracking target t, h i is the HOG feature of the i-th candidate box, and h t is the HOG feature of the tracking target; c t is the center point of the tracking box in the previous frame, ci is the center point of the detection box of the current frame. Finally, the candidate box with the largest total similarity score is selected as the real tracking object for tracking.
[0058] After the fusion features, the algorithm improves the ability to represent template objects. In addition to the apparent color features, the support of contour features can ensure that the tracker can make correct judgments when there is heavy traffic.
[0059] Step 4.4: Select the maximum score value after sorting in step 4.3, and combine it with the result calculated in step 3. When the IoU between the tracking box and the detection box is the largest and the Euclidean distance between their center points is the shortest, select the corresponding target as the appearance position of the tracking target in the current frame.
[0060] Step 5: When the tracking target is blocked, an anti-blocking algorithm based on vehicle motion state estimation is used. After tracking starts for more than 20 frames, the average speed of the vehicle in the video is recorded every 20 frames. According to this method, when the target disappears from the field of view, the coordinates at the time of disappearance are saved, and the target's moving speed is stopped and saved in the previous 20 frames. If the target disappears within 50 frames (about 3 seconds), the moving trajectory and coordinates within the disappearance are estimated normally, and the K neighborhood of the current estimated position is obtained at the same time, and the position where the target may appear after 3 seconds is recorded. The K neighborhood is set at this position to wait for the target to be captured. If the vehicle is re-detected and captured by the K neighborhood, the image matching algorithm in step 4 is called to start matching. If the match is successful, the tracking continues. If the match is not successful, the full image matching is turned on, the motion speed and coordinates are released, and the tracker searches by itself. The tracking method of steps 3-4 is called, and the multi-feature fusion matching algorithm is used to match the target with the greatest similarity to the initially selected tracking target in the current field of view, and the tracking frame is updated, and the K++ neighborhood is re-established based on the coordinates of the target.
[0061] Step 6: Determine whether the video has ended; if so, end the detection; otherwise, receive the next frame and return to step 2.
[0062] like Figure 3 As shown, Figure (3a) is the image of the 370th frame, which is the tracking effect in the unobstructed case. Figures (3b) and (3c) are the images of the 400th frame and 420th frame respectively, which are the tracking conditions when the traffic light partially obstructs the vehicle. The tracking algorithm of this embodiment can handle this situation well.
[0063] like Figure 4As shown, where Figures (4a), (4b), and (4c) are the images of the 146th frame, 219th frame, and 241st frame respectively. It can be seen that the image matching algorithm of this embodiment combines the K++ neighborhood search algorithm. Even when vehicles with similar colors and shapes appear near the target vehicle, the tracking algorithm can still handle it well.
[0064] As Figure 5 shown, where Figures (5a), (5b), and (5c) are the images of the 134th frame, 140th frame, and 144th frame respectively. The 134th frame indicates that the vehicle is about to disappear. The 140th frame is the motion estimation of the AOE anti-occlusion algorithm when the vehicle is completely occluded. The 144th frame is when the vehicle drives out of the occluder. At this time, the K++ neighborhood generated by the estimated box captures the target detected again and matches it with the tracking target. It can be seen that the AOE anti-occlusion algorithm can solve the problem of full occlusion of the target to a certain extent.
[0065] Using the method of this embodiment to perform single-target vehicle tracking and detection on urban roads, highways, and traffic congestion sections, the tracking accuracy is shown in Table 1. The average accuracy of the single-target vehicle tracking algorithm based on UAV aerial photography in this embodiment for vehicle tracking is 91.1%.
[0066] Table 1 Tracking accuracy of each scenario
[0067] Scenario Tracking accuracy Urban road 0.887 Highway 0.935 Traffic congestion 0.912
[0068] Compare the accuracy of the method of this embodiment with the accuracy of the TLD, SiamFC, SiamRPN++, and Dimp tracking algorithms. The results are shown in Table 2.
[0069] Table 2 Comparison of tracking accuracy of each algorithm
[0070] Algorithm Tracking accuracy The present invention 0.911 TLD 0.735 SiamFC 0.814 SiamRPN++ 0.876 Dimp 0.925
[0071] According to Table 2, compared with the traditional TLD algorithm and the single-target tracking algorithm based on the Siamese network, the accuracy of the target tracking algorithm of the present invention is affected by the detection results. Therefore, it is comprehensively evaluated that the tracking accuracy of the tracking algorithm in this article has increased by 5.9% on average.
[0072] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A single-target vehicle tracking method based on UAV aerial photography, characterized in that: The method includes the following steps: Step 1: Load the aerial video of the drone for vehicle tracking, pause at the first frame, manually select the target vehicle to be tracked with the mouse, and the selected area is the area to be tracked. The target vehicle within the area to be tracked is the tracking target. Then, starting from the second frame of the video, proceed to Step 2 to detect the target vehicle that appears in the video frame. Step 2: Determine whether the target vehicle in the current frame is fully occluded. If not, continue to Step 3; otherwise, execute Step 5. Step 3: Apply the K++ neighborhood search algorithm to predict the position of the tracking target. Establish a K++ neighborhood around the tracking box, use this K++ neighborhood to filter redundant target detection boxes, leave the vehicles that may be the tracking target appearing within the K++ neighborhood, and then calculate the IoU between the tracking box and the detection box and the offset of the center point. Step 4: Extract the target within the filtered detection box as a picture, and use the multi-feature fusion image matching algorithm to match it with the tracking target selected in the first frame. Calculate the image similarity between the tracking target and the filtered target respectively and sort them. Then, combine the calculation results of Step 3 to comprehensively determine which detected target in the current frame is the target to be tracked, and then update the tracking box and execute Step 6. The specific method is as follows: Step 4.1: After selecting the target to be tracked, extract the color histogram features and HOG features of the tracking target, and convert the two features into feature vectors. Step 4.2: In subsequent frames, extract the detected targets of the same category as pictures, and also extract the color histogram features and HOG features of each target to obtain feature vectors. In the calculation of the color histogram feature vectors in Step 4.1 and Step 4.2, divide the value range of 0-255 of each primary color into four equally sized regions: [0,63] is region 0, [64,127] is region 1, [128,191] is region 2, and [192,255] is region 3; each primary color value has four corresponding values after partitioning. Any color that appears in the image will definitely belong to one of these four regions. Count the number of pixels that appear in each region to obtain a 64-dimensional feature vector. Step 4.3: Calculate the color histogram feature similarity and HOG feature similarity between the tracking target and all the targets obtained in Step 4.2 respectively, and perform weighted scoring. Finally, sort all the scores; finally, select the candidate box with the highest total similarity score as the real tracking object for tracking. The calculation method of the color histogram feature similarity is cosine similarity, which satisfies Equation (5): The smaller the included angle, the larger the cosine value corresponding to the included angle, which means they are more similar; P and Q are two multi-dimensional vectors, where P is [P1, P2,..., P m , and Q is [Q1, Q2,..., Q m , and m is the dimension of the vector; The HOG feature similarity uses the HOG feature descriptor, that is, the feature vector, and finally calculates the Euclidean distance between the feature vectors. The smaller the distance, the more similar the two pictures are; since the HOG feature vector is an n-dimensional vector, the corresponding Euclidean distance satisfies Equation (6): where d is the Euclidean distance of the vector, and x i and y i are two coordinate values of the vectors in the multi-dimensional space; After calculating the cosine similarity of the color histogram features and the Euclidean distance similarity of the HOG features, multiply them by their corresponding weights respectively and add them to obtain the final similarity score, which satisfies Equation (7): S i = W1(S(c i , c t )) + W2(S(h i , h t )) (7) Where, S i is the total similarity score between the image in the i-th candidate box and the tracking target; W1 is the similarity weight coefficient of the color histogram feature; W2 is the similarity weight coefficient of the HOG feature; S is the similarity calculation function of the parameters in the brackets; S(c i , c t ) is the color histogram feature similarity function between the i-th candidate box and the tracking target t, c i is the center point of the detection box in the current frame, c t is the center point of the tracking box in the previous frame; S(h i , h t ) is the HOG feature similarity function between the i-th candidate box and the tracking target t, h i is the HOG feature of the i-th candidate box, h t is the HOG feature of the tracking target; Step 4.4: Select the maximum score value after sorting in step 4.3, and combine it with the result calculated in step 3. When the IoU between the tracking box and the detection box is the largest and the Euclidean distance between their center points is the shortest, select the corresponding target as the appearance position of the tracking target in the current frame; Step 5: When the tracking target is occluded, an anti-occlusion algorithm based on vehicle motion state estimation is used. After the tracking starts for more than 20 frames, the average speed of the vehicle in the video is recorded every 20 frames. According to this method, when the target disappears from the field of view, the coordinates at the time of disappearance are saved, and the target's moving speed is stopped and saved in the previous 20 frames. If the target disappears within 50 frames, the moving trajectory and coordinates within the disappearance are estimated normally, and the K neighborhood of the current estimated position is obtained at the same time, and the position where the target may appear after 50 frames is recorded. The K neighborhood is set at this position to wait for the target to be captured. If the vehicle is re-detected and captured by the K neighborhood, the image matching algorithm in step 4 is called to start matching. If the match is successful, tracking continues. If the match is not successful, full-image matching is turned on, and the recording of the moving speed and coordinates is released. The tracker searches for it by itself, and the tracking method of steps 3-4 is called. The multi-feature fusion matching algorithm is used to match the target with the greatest similarity to the initially selected tracking target that appears in the current field of view, and the K++ neighborhood is re-established based on the coordinates of the target; execute step 6; Step 6: Determine whether the video has ended; if so, end the detection; otherwise, receive the next frame and return to step 2.
2. The single-target vehicle tracking method based on UAV aerial photography according to claim 1, wherein: When selecting a target in step 1, only the target itself is selected.
3. The single-target vehicle tracking method based on UAV aerial photography according to claim 1, characterized in that: The specific method of step 3 is: Step 3.1: According to the size of the tracking frame of the previous frame, calculate the K neighborhood range corresponding to the tracking frame when k=2, and narrow the detection range of the current frame to the K neighborhood; the K neighborhood satisfies formula (1): Among them, W k and H k are respectively the width and height of the K-neighborhood search area, W and H are respectively the width and height of the target tracking box in the previous frame, and k is the aspect ratio of the width to the height of the two; Step 3.2: If there is only one target detected in the K neighborhood in the current frame, that is, at least two-thirds of the target's detection frame is within the K neighborhood, then the target is the target of the previous frame, update the tracking frame, and continue to step 4; if there are more than two targets in the K neighborhood in the current frame, execute step 3.3; Step 3.3: Calculate the similarity between the targets in the K neighborhood and the tracked target, obtain the similarity scores, and sort them; Step 3.4: Compare the target detection frame corresponding to the sorted similarity score with the tracking frame of the previous frame and the Euclidean distance of the center point; take the detection frame corresponding to the center point with the smallest Euclidean distance, and combine it with the similarity calculated by the previous image matching and the largest IoU to determine which detection frame detects the tracking target. The judgment order is: first compare the image similarity, then exclude similar vehicles based on IoU, and finally use the Euclidean distance of the center point to select the tracking target.
4. The single-target vehicle tracking method based on UAV aerial photography according to claim 3, wherein: In step 3.4, the calculation of IoU satisfies formula (2): Among them, gt is the tracking box of the previous frame; bb is the detection box that appears within the K-neighborhood range in the current frame. Calculate the IoU using gt and the detection boxes within the K-neighborhood respectively, and select the detection box with the largest IoU value for retention, satisfying Equation (3): IoU(gt,bb) max = Max(IoU(gt,bb1),...,IoU(gt,bb n )) (3) Among them, IoU() is the intersection over union between the detection box and the tracking box of the previous frame, gt is the tracking box of the previous frame, and bb i is the i-th detection box that appears in the K++ neighborhood of the current frame, and n is the total number of detection boxes in the K++ neighborhood; Calculate that the Euclidean distance of the center point satisfies Equation (4): Among them, d(·) is the Euclidean distance between two points in the parentheses, c1 and c2 are the center points of two boxes, and their coordinates are (x1, y1) and (x2, y2) respectively; c gt is the center point of the tracking box in the previous frame, and c bb is the center point of the detection box in the current frame.
Citation Information
Patent Citations
Unmanned aerial vehicle single-vision target tracking method for target shielding
CN111862155A
Pedestrian tracking method based on YOLOv3
CN112884810A