Unmanned aerial vehicle target tracking method based on local image data learning

By adopting the discriminant tracking method of feature and structure support vector machine classifier based on local image data learning in drone videos, combined with significance information, the problem of poor target tracking effect in drone videos is solved, and the tracking effect with high accuracy and success rate is achieved.

CN119941786APending Publication Date: 2025-05-06SHANGHAI LIONWEI INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411867435.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2018-05-14
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing video target tracking methods are not effective in drone videos, especially when the target is small, the distance is long and the scene changes greatly, it is difficult to track effectively.

Method used

The features based on local image data learning are adopted, combined with the structural support vector machine classifier, a discriminant tracking method is adopted, and the significance information of the target is used for auxiliary tracking.

Benefits of technology

It achieves high accuracy and success rate for targets in drone videos, especially in complex scenarios such as occlusion, light changes, and scale changes, which can maintain good tracking effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941786A_ABST
    Figure CN119941786A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle target tracking method based on local image data learning, and the method comprises the following steps: 1, carrying out the sampling of a position near a previous frame of a target through employing a particle filtering method, and obtaining a potential target; step 2, calculating features of each potential target according to the potential targets obtained in the step 1; 3, classifying the targets by using a structural support vector machine, and searching for an optimal target; 4, calculating the feature saliency around the object; 5, fusing the target position by using the saliency information obtained in the first step and the target position obtained in the third step, and determining a final target position; 6, sampling at the final position of the target to obtain positive and negative samples for training, and updating the model of the structural support vector machine online; the target positions are fused and calculated as follows: y = alpha1y1 + alpha2y2, y is the finally determined target, y1 and y2 are the target positions obtained in the step 3) and the step 4) respectively, and alpha1 and alpha2 are weights respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for implementing target tracking by utilizing computer pattern recognition technology, and in particular to a method for unmanned aerial vehicle target tracking based on local image data learning. Background Art

[0002] In traditional video target tracking, classic tracking algorithms such as subspace learning, sparse representation, multi-task learning, and multi-instance learning all use various features to model the target, and then classify the potential target by comparing it with the target or using the learned classifier to achieve the purpose of target tracking. They have achieved certain success on general target tracking databases. According to the tracking method, tracking methods can usually be divided into two types: generative and discriminative tracking methods.

[0003] The generative tracking method is to model the potential target, and then compare it with the existing target through some kind of metric. The closest one is considered to be the target. Earlier video target tracking used the color histogram of the target as a feature, and compared the potential target with the target template through the Bhattacharyya coefficient (Condensation—conditional density propagation for visual tracking, IJCV, 29 (1), 5-28, 1998). The subspace method performs PCA decomposition on the target and obtains the principal component of the target as a feature for tracking (Incremental learning for robust visual tracking, IJCV, 77, 125-141, 2008). Superpixel features are also used as target features in target tracking (Superpixel tracking, ICCV, 2011). After sparse representation was introduced into computer vision, it was also used for target tracking (Robust visual tracking using l1minimization, ICCV, 2009). Here, the pixel information of the target is directly used as the target feature, and the target is found by comparing the distance between the potential target and the template 1 norm. The paper Robust visual tracking via multi-task sparse learningin mentions the introduction of multi-task learning method into target tracking (IEEE conference on computer vision and pattern recognition, 1-8, 2012).

[0004] The above target tracking methods are all based on the idea of ​​minimizing errors, while the discriminant method uses the data in the image to train a classifier, and uses the classifier to classify the targets in the image, thereby finding the foreground target. The multiple instance learning algorithm calculates the Harr features of the target, and then uses the multiple instance algorithm to classify the potential targets (Vi suall tracking with online multiple instance learning, in CVPR, 2009). The literature (Support vector tracking, TPAMI, 26 (8), 1064-1072, 2004) uses the support vector machine (Support Vector Machine, SVM) as a classifier. The literature (Structured output tracking with kernels. In ICCV, 263-270, 2011) uses the structure support vector machine (Structure Support Vector Machine, SSVM) for classification. The use of the structure support vector machine classifier has achieved good results in target tracking experiments due to its fast speed and high accuracy.

[0005] However, most of the currently known tracking methods are for tracking general targets, such as people and cars. In ordinary videos, the targets of people and cars are relatively large and are viewed horizontally. Such images have been studied in computer vision for a long time, so the existing features can describe these objects well. In target tracking, features have a great influence on the tracking effect. For example, in the above tracking methods, image histograms, pixel intensity, Haar features, and superpixel features are used. For the targets in drone tracking, because the distance is far, the targets are very small and only the outline of the targets can be seen. Features are needed to better obtain the edge information of the targets. In addition, the target scenes vary greatly, which requires the features to have a certain learning ability. Summary of the invention

[0006] In order to solve the defects of the prior art, the present invention proposes a feature based on local data learning, combined with a structured support vector machine classifier, adopts a discriminant tracking method, and uses the saliency information of the target to assist tracking.

[0007] A target tracking method based on local image data learning includes learning based on local image data to obtain local features, and then implementing target tracking in combination with a fusion video of significant features under the framework of discriminant tracking.

[0008] Another target tracking method based on local image data learning includes the following steps: (1) dense sampling is performed around the target, and the sampled targets are potential targets, and those outside the sampling frame are discarded; (2) local features of the image are extracted based on local data, and local data are learned at the same time to improve the stability of target changes and the robustness to image noise and geometric deformation; (3) a structural support vector machine is used as a classifier, and a discriminant tracking method is adopted to improve the classification ability of the extracted local features, while overcoming the problem of tracking loss caused by poor feature selection in the previous structural support vector machine classifier; (4) the saliency features of the image are calculated to obtain the position of the object; (5) the tracking position of the structural support vector machine and the object position obtained by saliency are fused to determine the final position of the object. (6) According to the prediction, the positive and negative support vectors are updated.

[0009] Therefore, the local features selected by this method have good classification ability and are not very sensitive to changes in lighting conditions, posture, etc., and are robust to slight physical noise and geometric deformation.

[0010] Another technical solution provided by the present invention is to realize video tracking of drones by combining with a structural support vector machine, thereby achieving a higher accuracy and success rate in drone tracking experiments.

[0011] A target tracking method based on local image data learning is applied in UAV video tracking. The method comprises the following steps:

[0012] The first step is to use the particle filter method to sample the vicinity of the target's previous frame position to obtain potential targets.

[0013] In the second step, the features of each potential target are calculated based on the potential targets obtained in the first step.

[0014] The third step is to use the structural support vector machine to classify the targets and find the best target.

[0015] The fourth step is to calculate the feature saliency around the object.

[0016] In the fifth step, the saliency information obtained in the first step and the target position obtained in the third step are used to fuse the target position and determine the final target position.

[0017] In the sixth step, sampling is performed at the final position of the target to obtain positive and negative samples for training, and the model of the structural support vector machine is updated online.

[0018] The present invention proposes a local feature with good ability to describe target edge information and a certain degree of robustness to noise and geometric deformation, designs a discriminant tracking method in combination with a structural support vector machine (SSVM), and uses the saliency information of the target to assist tracking. The tracking method proposed in the present invention can achieve better tracking results in various scenarios (such as occlusion, light changes, and scale changes). The present invention can be applied to various civil and military systems such as face tracking, drone tracking, and military target tracking systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of an embodiment of the present invention.

[0020] Figure 2 Schematic diagram of geodesic distance.

[0021] Figure 3 Schematic diagram of significant information.

[0022] Figure 4 Schematic diagram showing the experimental results of the accuracy of several methods on the drone database.

[0023] Figure 5 Schematic diagram of experimental results showing the success rates of several methods on the drone database.

[0024] Figure 6 Schematic diagram showing the experimental results of the accuracy of several methods on the drone database (scale change scenario).

[0025] Figure 7 Schematic diagram showing the experimental results of the accuracy of several methods on the drone database (scene with aspect ratio changes).

[0026] Figure 8 Schematic diagram showing the experimental results of the accuracy of several methods on the drone database (low-resolution scene).

[0027] Fig. 9 Schematic diagram of experimental results showing the success rate of several methods on the drone database (scale change scenario).

[0028] Fig.10 Schematic diagram of experimental results showing the success rate of several methods on the drone database (scene with aspect ratio changes).

[0029] Fig.11 Schematic diagram showing the experimental results of the success rates of several methods on the drone database (low-resolution scene).

[0030] Fig.12 It shows the tracking result of bicycle video on the UAV database by the method of the present invention.

[0031] Fig.13 It shows the tracking result of the method of the present invention in the video of embarking on the ship in the UAV database.

[0032] Fig.14 It shows the tracking result of the shipboard video in the UAV database by the method of the present invention.

[0033] Fig.15 It shows the tracking result of the crowd video on the drone database by the method of the present invention. DETAILED DESCRIPTION

[0034] The following is a detailed description of an embodiment of the present invention in conjunction with the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiments.

[0035] like Figure 1 As shown, the purpose of this embodiment is first to make the features of the present invention have a good ability to describe the target, and secondly to make the features able to overcome the problem of noise interference. The approach taken is to use the geodesic distance on the manifold instead of the Euclidean distance, and consider the gradient of the image. Classification is performed through SSVM to determine the target position. Then, the saliency information of the target is considered, and the target is finally located through fusion.

[0036] The technical solution of the present invention involves geodesic distance, local features, structural support vector machine method, and calculation of significance method, and is described in detail as follows.

[0037] The geodesic distance involved in the present invention is described as follows:

[0038] Figure 2 The difference between geodesic distance and Euclidean distance is given. Assuming that on a surface, the data can be represented as {x 1 ,x 2 ,z(x 1 ,x 2 )}., where x 1 and x 2 is the coordinate, z(x 1 ,x 2 ) indicates that the surface changes with x 1 and x 2 Change, that is, the surface can be expressed as S(x 1 ,x 2 )={x 1 ,x 2 ,z(x 1 ,x 2 )}. Then the length of the curve in the figure can be expressed as:

[0039] The local features of the present invention are described as follows:

[0040] If we expand the above formula for calculating the length of the curve, we can get:

[0041]

[0042] Where Δx=[dx 1 , dx 2 ] T , Since we use a local operator here, +Δx T Δx is a very small variable and can be ignored, so ds 2 ≈Δx T CΔx. Therefore, our local feature can be expressed as: K(C, Δx) = exp(-ds 2 )=exp{-Δx T CΔx}, where C is calculated as follows:

[0043]

[0044] In the formula, u 1 and u 2 represents the eigenvector of matrix C, s 1 and 2 The values ​​of ∈, τ, and α are chosen to be 10 respectively. -8 , 0.4 and 0.2.

[0045] The structural support vector machine method involved in the present invention is described as follows:

[0046] It is also used to deal with multi-classification problems. The advantage of support vector machines is that they have a good theoretical basis, that is, they have strong generalization ability. Its disadvantages are 1) high training complexity; 2) they cannot be used to predict structured problems. Structural support vector machines extend support vector machines from binary classification problems to predicting structured problems by modifying the constraints and objective functions of support vector machines.

[0047] The goal of structural support vector machine training is to learn the prediction function f. In the following formula, g is the scoring function. The scoring function outputs a continuous value instead of the traditional [+1,-1]. y is a rectangle in the search space, t is the current frame, time space, and function f is used to map time space to coordinate space, that is, the function of predicting the position process

[0048] f(t)=argmax y∈Y g(t,y)

[0049] Using a series of sample training to solve the following problem is the loss function of the basic support vector machine, which belongs to the maximum margin framework, minimizing w, with a soft margin term:

[0050]

[0051]

[0052] The Δ symbol represents the interval we defined. The actual meaning of this interval is related to the overlap rate. When the overlap rate is 1, the value is 0, which is equivalent to the sample at the center during training, and the interval is set to 0.

[0053]

[0054] Represents the two tracking positions y and The degree of overlap.

[0055] The calculation of the significant features involved in the present invention is described as follows:

[0056] Image saliency is an important visual feature in an image, which reflects the degree of attention that the human eye pays to certain areas of the image. Since the work of Itti (L. Itti, C. Koch, & E. Niebur. A model of saliency based visual attention for rapid scene analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(11): 1254-1259, 1998.) in 1998, a large number of saliency mapping methods have been developed, and image saliency has also been widely used in image compression, coding, image edge and region enhancement, salient object segmentation and extraction, etc. For an image, the user is only interested in some areas of the image. These areas of interest represent the user's query intent, while most of the remaining uninterested areas are irrelevant to the user's query intent. The salient area is the area in the image that can most arouse the user's interest and best express the image content. In fact, the selection of salient areas is very subjective. Due to different user tasks and knowledge backgrounds, different users may choose different areas as salient areas for the same image. The commonly used method is to calculate the saliency of the image based on the human attention mechanism. Research in cognitive psychology shows that some areas in an image can significantly attract people's attention, and these areas contain a large amount of information. Cognitive scientists have proposed many mathematical models to simulate people's attention mechanisms. Since the general rules of the image cognitive process are utilized, the extracted salient areas are more consistent with people's subjective evaluation. In this patent, we use image saliency to infer the location of the target.

[0057] After obtaining the local eigenvalues ​​of the image, we use the matrix cosine similarity measure (MCS) to calculate the similarity between the features at each position and the features at other positions. The saliency is calculated as follows:

[0058]

[0059] Among them, F i and F j Represent the local features at positions i and j respectively.

[0060] The present invention fuses the tracking position of the structural support vector machine and the object position obtained by saliency to determine the final object position. The calculation method for fusion of the target position is as follows:

[0061] y=α 1 y 1 +α 2 y2

[0062] Among them, y is the final target, y 1 and 2 are the position obtained by the structural support vector machine and the target position obtained by the salient features, α 1 and α 2 are the weights respectively.

[0063] On the basis of the above technologies, the technical solution of the present invention is specifically implemented as follows:

[0064] 1) Sampling to obtain potential targets

[0065] Dense sampling is performed around the target, and the sampled objects are potential targets, and those outside the sampling frame will be discarded.

[0066] 2) Calculate local features.

[0067] For each potential object in the image, we compute local features.

[0068] 3) Use structural support vector machine classification to obtain the target location.

[0069] Use a support vector machine to output a score for the location of each potential target and predict the location of the target in the current frame.

[0070] 4) Calculate significance.

[0071] A box area is determined near the target in the previous frame obtained by tracking, and the saliency of the image is calculated.

[0072] 5) Position fusion.

[0073] The target positions obtained in step 3) and step 4) are fused and calculated as follows:

[0074] y=α 1 y 1 +α 2 y 2 ,

[0075] Among them, y is the final target, y 1 and 2 are the target positions obtained in step 3) and step 4), α 1 and α 2 are the weights respectively.

[0076] 6) Online update.

[0077] Based on the predictions, the positive and negative support vectors are updated, and then the structural support vector machine model is updated.

[0078] The experimental data uses the UAV database from: https: / / ivul.kaust.edu.sa / Pages / Dataset-UAV123.aspx. This source provides 123 videos and 12 scenes. The longest video has 3085 frames and the total number of frames exceeds 110,000. The tracking scenes are quite extensive, including: roads, near buildings, and grass. The tracked targets include: people, cars, and ships. The 12 scenes are:

[0079] Aspect ratio change: The aspect ratio of the object changes beyond the range of [0.5,2].

[0080] Background Clutter: The background and the target are very similar.

[0081] Camera Motion: The camera shakes suddenly.

[0082] Fast Motion: the object moves very quickly between the previous and next frames.

[0083] Full Occlusion: The target is completely occluded.

[0084] Illumination Variation: The brightness of the light on the surface of an object changes dramatically.

[0085] Low Resolution: the size of the object is less than 400 pixels.

[0086] Out-of-View: Part of the object is out of the camera range.

[0087] Partial Occlusion: Part of the object is occluded.

[0088] Similar objects (SimilarObject), there are other similar objects near the target.

[0089] Scale Variation: The scale variation of the object exceeds the range of [0.5,2].

[0090] Viewpoint Change: The angle at which you observe an object changes significantly.

[0091] At the same time, the algorithms compared with the method of this embodiment include: kernel correlation filtering tracking method (shown as "Comparison Algorithm 1" in the figure, which is a sparser dotted line), detection learning tracking algorithm (shown as "Comparison Algorithm 2" in the figure, which is a denser dotted line) and support vector machine-based tracking algorithm (shown as "Comparison Algorithm 3" in the figure, which is a long and short dotted line).

[0092] Two commonly used evaluation criteria in computer vision are used: precision score and success score. Precision score indicates the difference between the center of the tracked position and the center of the true value position. Success score indicates the overlap rate between the tracked position box and the true value position box.

[0093] Figure 4 Indicates the accuracy of different methods on the drone database. In the figure, "our algorithm" represents the tracking method of the present invention, which is represented by a solid line. It can be seen that the present invention has achieved a better accuracy on the database drone.

[0094] Figure 5 It represents the success rate of different methods on the UAV database. It can be seen that the present invention has achieved a better success rate on the database UAV.

[0095] Figure 6-Figure 8 It shows the accuracy of different methods in various scenarios on the drone database, with a total of 12 scenarios. It can be seen that the present invention has achieved a better accuracy in each scenario.

[0096] Figure 9-11 It represents the success rate of different methods in various scenarios on the drone database. It can be seen that the present invention has achieved a good success rate in various scenarios.

[0097] Fig.12 The figure shows the tracking result of the bicycle video on the UAV database using the method of the present invention.

[0098] Fig.13 It shows the tracking result of the method of the present invention on the ship video in the UAV database.

[0099] Fig.14 It shows the tracking result of the shipboard video on the UAV database by the method of the present invention.

[0100] Fig.15 The figure shows the tracking results of the method of the present invention on the crowd video on the drone database.

Claims

1. A UAV target tracking method based on local image data learning, characterized in that The steps include: The first step is to use the particle filter method to sample the vicinity of the target's previous frame position to obtain potential targets; In the second step, based on the potential targets obtained in the first step, the features of each potential target are calculated; The third step is to use the structural support vector machine to classify the targets and find the best target; The fourth step is to calculate the feature saliency around the object; Step 5: Use the saliency information obtained in the first step and the target position obtained in the third step to fuse the target position and determine the final target position. Step 6: Sample at the final position of the target to obtain positive and negative samples for training, and update the model of the structural support vector machine online; The target positions are fused and calculated as follows: y=a1y1+a2y2, Among them, y is the final target, y1 and y2 are the target positions obtained in step 3) and step 4), and α1 and α2 are weights.

2. The method for tracking unmanned aerial vehicle targets based on local image data learning according to claim 1, characterized in that: The local features are expressed as: K(C,Δx)=xp(-ds 2 )=exp{-Δx T CΔx}, Among them, C is calculated as follows: Where u1 and u2 represent the eigenvectors of matrix C, s1 and s2 represent the eigenvalues ​​of matrix C, and the values ​​of ∈, τ, and α are selected as 10 -8 , 0.4 and 0.

2.

3. The method for tracking unmanned aerial vehicle targets based on local image data learning according to claim 1, characterized in that: The goal of the step (3) structural support vector machine training is to learn the prediction function f, where g is the scoring function. The scoring function outputs a continuous value instead of the traditional [+1, -1]. y is a rectangle in the search space, t is the current frame, time space, and function f is used to map time space to coordinate space, that is, the function of predicting the position process f(t)=argmax y∈Y g(t,y); Using a series of sample training to solve the following problem is the loss function of the basic support vector machine, which belongs to the maximum margin framework, minimizing w, with a soft margin term: The Δ symbol represents the interval we defined. The actual meaning of this interval is related to the overlap rate. When the overlap rate is 1, the value is 0, which is equivalent to the sample at the center during training, and the interval is set to 0. Represents the two tracking positions y and The degree of overlap.

4. The method for tracking unmanned aerial vehicle targets based on local image data learning according to claim 1, characterized in that: The saliency of the computer image is calculated as follows: Among them, F i and F j Represent the local features at positions i and j respectively.