Unmanned aerial vehicle target tracking method based on target deformation and feature fusion

By combining the tracking anomaly detection network with target deformation trend awareness vectors and feature fusion, the filter is dynamically updated, and the problem of insufficient robustness of the drone target tracking algorithm when dealing with abnormal situations is solved, achieving stronger tracking performance and computing efficiency.

CN119941788APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411950443.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing drone target tracking algorithms are not robust enough to deal with abnormal situations such as target deformation, occlusion, perspective transformation and fast motion, and the calculation speed based on deep learning methods is slow, making it difficult to meet the hardware requirements of the drone platform.

Method used

The drone target tracking method based on the fusion of target deformation and feature is adopted, and filtering is performed through fast Fourier transform, combined with the target deformation trend-aware vector and the tracking anomaly detection network of feature fusion, the filter is dynamically updated to improve tracking performance.

Benefits of technology

Effectively dealing with dynamic learning and exceptions of the target improves the robustness of the tracking algorithm, avoids the introduction of background noise, and improves the tracking performance while ensuring computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941788A_ABST
    Figure CN119941788A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle target tracking method based on target deformation and feature fusion, and the method comprises the steps that a target deformation trend perception module carries out the second-order difference of a response diagram among three continuous frames, obtains the change acceleration of each pixel, and achieves the perception of a target deformation trend; the self-adaptive space-time regularization module dynamically modifies a target regularization mask in a spatial domain by sensing a target deformation trend, and dynamically modifies a target time domain penalty term coefficient in a time domain, so that construction of a self-adaptive space-time regularization term is completed; and the feature fusion-based anomaly detection module obtains the feature fusion-based anomaly detection module by performing fusion training on the features of the response diagram and the target deformation trend sensing data when the tracking anomaly occurs by using the distribution distortion in the response diagram and the target deformation trend sensing data. The efficient unmanned aerial vehicle target tracking method designed by the invention can effectively improve the tracking performance of a tracker in various complex environments, and has a certain promotion effect on further promoting target tracking application based on an unmanned aerial vehicle platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion. Background Art

[0002] Vision-based target tracking is one of the basic tasks of computer vision. Its goal is to locate the target in the subsequent continuous image sequence based on the position of the target in the first frame of the image. With the rapid development and maturity of drone technology in recent years, target tracking based on drone platforms has been more widely used, such as tracking aerial photography, pedestrian following, security patrol, etc., and has broad development prospects in the military and civilian fields.

[0003] Since the UAV platform itself is flexible and has a changeable perspective, it can adapt to tracking tasks in a variety of different scenarios. However, such flexibility is both an advantage of the UAV platform and a challenge to the tracking algorithm. In real tasks, it is likely to encounter tracking anomalies such as target deformation, target occlusion, perspective change, rapid movement, and out of field of view, so higher requirements are placed on the robustness of the tracking algorithm. And considering that the computing power and energy carried by the UAV platform are limited, the tracking algorithm must ensure good tracking performance while reducing computing energy consumption as much as possible. In this regard, there are currently two main research directions, one is the target tracking method based on the correlation filtering algorithm, and the other is the target tracking method based on the deep learning algorithm. The target tracking method based on the correlation filtering algorithm can generally achieve a smaller computational overhead because it can convert complex convolution operations into the frequency domain for processing, but it is limited by the algorithm's limited interpretation of image depth information, and it is difficult to achieve the tracking accuracy of the target tracking method based on the deep learning algorithm. However, since the target tracking methods based on deep learning algorithms are generally based on multi-layer convolution, a large amount of data training is required to achieve good tracking performance, and a lot of time is needed in both the training and reasoning stages. In addition, considering the limited computing power resources of the UAV airborne platform, it is difficult to support the deployment of the existing target tracking methods based on deep learning algorithms with good tracking effects. Therefore, the method of using the target tracking method based on the deep learning algorithm to ensure the tracking performance of the tracker is usually slow in calculation speed and cannot meet the hardware requirements for deployment and the implementation requirements of tracking. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a UAV target tracking method based on target deformation and feature fusion, which overcomes the shortcomings of the current target tracking method based on correlation filtering in dynamic learning of the target and handling of abnormal situations.

[0005] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:

[0006] A method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion includes the following steps:

[0007] 1) Obtain the target image through the drone’s onboard camera;

[0008] 2) Construct a tracking algorithm model and use the model to calculate the position of the target to be tracked in the image;

[0009] 3) Update the filter in the tracking algorithm model and wait for the next frame of image input.

[0010] The step 2) comprises the following steps:

[0011] 2.1) The current frame image and the filter are transformed into the frequency domain by the fast Fourier algorithm and then filtered to obtain a response graph;

[0012] 2.2) The peak position of the response graph is used as the position of the target to be tracked in the current frame.

[0013] The step 3) comprises the following steps:

[0014] 3.1) The response map is stored in the response map vector, and the target deformation trend perception vector is calculated;

[0015] 3.2) Calculate the target spatial domain mask and temporal domain regularization term coefficient according to the target deformation trend perception vector;

[0016] 3.3) Input the response graph and the target deformation trend perception vector into the tracking anomaly detection network based on feature fusion, and judge whether the current tracking state is normal according to the detection results;

[0017] 3.4) If there is no tracking anomaly, the spatial domain mask and the temporal domain regularization term coefficients are updated to the objective function, a new filter is calculated by the ADMM algorithm, and the original filter is updated; if there is a tracking anomaly, the original filter is kept unchanged.

[0018] The calculation target deformation trend perception vector is specifically:

[0019] The second-order response diagram is taken as the rate of change vector of the response diagram Λ=[|Λ 1 |,|Λ 2 |,|Λ 3 |,…,|Λ T |], the i-th element |Λ i |For:

[0020]

[0021]

[0022] in, represents the i-th element of the t-th frame response graph, Φ Δ represents an alignment shift operation, represents the change of the i-th element of the response graph of the t-th frame, Λ i Indicates the change rate corresponding to the current response graph, that is, the target deformation trend perception vector.

[0023] The target spatial domain mask W is specifically:

[0024] W=w+f

[0025] Among them, w is the basic Gaussian mask:

[0026]

[0027] σ=std[δlog(Λ+1)]

[0028] Among them, μ X and μ Y represents the mean of the two coordinates, σ represents the standard deviation of the normalized deformation perception Λ, δ is the scaling factor, std represents the standard deviation calculation function, X, Y represent the horizontal and vertical coordinates;

[0029] f is the local refined Gaussian mask:

[0030] f(X,Y)=max(f1,f2,…,f K )

[0031]

[0032] Among them, k represents the index of each category, and each two-dimensional Gaussian distribution is optimally obtained to obtain a local refined Gaussian mask f.

[0033] The time domain regularization term coefficient θ t Specifically:

[0034]

[0035] in, is the time domain regularization coefficient θ t , ReLU represents the ReLU function, and ζ and ν are hyperparameters.

[0036] The step 3.3) comprises the following steps:

[0037] 3.3.1) The single-channel response map and the single-channel change rate matrix of size H×W are respectively passed through a convolutional layer with the same output dimension to obtain a C×H×W feature map;

[0038] 3.3.2) Add and fuse the two feature maps;

[0039] 3.3.3) Learning channel features and global features through channel feature branches and global feature branches respectively;

[0040] 3.3.4) The feature maps output by the two branches are added and fused, and then mapped to the range of 0 to 1 through the Sigmoid activation function, which is used as the weight of the response map distribution;

[0041] 3.3.5) After the single channel response map and the single channel change rate matrix are weighted and added together according to the weights, the deep feature extraction is further performed through the convolution layer;

[0042] 3.3.6) Use the weighted features as the input of the activation function, output the feature value, and convert it into a probability distribution through the fully connected layer Softmax;

[0043] 3.3.7) The classification of the current response graph is obtained according to the probability of each category, and the judgment of whether an abnormality occurs during the tracking process is completed.

[0044] The channel feature branch performs the following steps:

[0045] The channel features are passed through the pixel convolution layer to obtain the deep features of C1×H×W, and then after passing through the ReLU activation function, another pixel convolution layer restores the feature dimension to C×H×W.

[0046] The global feature branch performs the following steps:

[0047] After the global features are pooled through global average, the deep features of C1×H×W are obtained by the pixel convolution layer. After the ReLU activation function, the feature dimensions are restored to C×H×W by another pixel convolution layer.

[0048] The present invention has the following beneficial effects and advantages:

[0049] The present invention can effectively address the deficiencies of target tracking methods in dynamic learning of targets and handling of abnormal situations, can timely update the model according to target deformation and prevent the introduction of background noise in abnormal tracking conditions, and ultimately achieve tracking performance with stronger robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a structural diagram of the UAV target tracking model based on target deformation and feature fusion of the present invention. DETAILED DESCRIPTION

[0051] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0052] like Figure 1 As shown, the present invention includes the following processes:

[0053] First, the image of the RGB channel is input into the model; the image and the original filter are converted to the frequency domain, and the response map of the frame image is calculated by the correlation filtering algorithm. Then, the current position of the target is obtained and output through the peak position of the response map; the current target deformation trend perception vector is calculated through the response map, and the target spatial domain mask and time domain regularization term coefficient are dynamically corrected according to the vector; the response map and the target deformation trend perception vector are input into the tracking anomaly detection network based on feature fusion to determine the current tracking state; when the tracking state is normal, the target spatial domain mask and time domain regularization term coefficient are updated to the objective function, and the filter is updated through the ADMM algorithm; when the tracking state is abnormal, the original filter is retained; the filter is updated or retained after the filter is updated, and the next frame of image input is waited for.

[0054] The specific steps are as follows:

[0055] Step 1. Input image data.

[0056] The image data to be tracked is input into the UAV target tracking model based on target deformation and feature fusion.

[0057] The image is a current RGB image or a single-channel image read from the drone's onboard camera.

[0058] Step 2. Calculate the response map of the frame image through the correlation filtering algorithm based on the current read image and the original filter.

[0059] The input image and the obtained filter are transformed into the frequency domain through the FFT algorithm, and then the correlation filtering calculation is performed to obtain the response graph. In the second frame, the filter is the initialization filter of the first frame, and the filter used afterwards is the new filter updated after learning in the previous frame.

[0060] Step 3. Confirm the position of the target to be tracked in the current frame based on the peak position of the response graph.

[0061] The position of the peak of the response graph is the position of the target in the current frame, that is, the output of the tracking algorithm.

[0062] Step 4. Store the response map into the response map vector and calculate the target deformation trend perception vector.

[0063] The change of the response graph reflects the change of the target appearance and scene, and its change trend can reflect the change of the target's motion state and scene illumination. Based on this, this study further explores the deep information in the response graph, and introduces the second-order response graph as the change rate vector of the response graph Λ=[|Λ 1 |,|Λ 2 |,|Λ 3 |,…,|Λ T |], the i-th element |Λ i|The definitions are as follows:

[0064]

[0065] in, represents the i-th element of the t-th frame response graph, Φ Δ Indicates the alignment shift operation, aligning the peak positions of the response graphs of the previous and next two frames. represents the change of the i-th element of the response graph of the t-th frame, Λ i Indicates the change rate corresponding to the current response graph, that is, the target deformation trend perception vector.

[0066] When the target shape and scene environment remain unchanged or change relatively smoothly, the response graph change rate Λ will remain low because the change amount between the two frames is very close. However, when the target motion state or external scene changes, such as when the target turns, is blocked, the illumination changes, or leaves the field of view, the change rate will cause Λ to change greatly. In other words, the value of the change rate Λ is positively correlated with the severity of the target change.

[0067] Step 5. Calculate the target spatial domain mask and temporal domain regularization term coefficient based on the target deformation trend perception vector.

[0068] Adaptive spatial regularization: Spatial regularization is a quantitative representation of the credibility of each pixel. It adds different weights to pixels at different locations by adding masks, thereby changing the learning of the model in different areas. Generally speaking, under normal tracking conditions, the mask is a Gaussian distribution centered on the target. Thanks to the target deformation perception Λ proposed in this study, this study can not only achieve adaptive adjustment of the basic Gaussian mask, but also further refine the weight distribution of parts with larger deformation trends on the basic Gaussian mask, and further improve the learning ability of the model when the target is deformed by optimizing the mask.

[0069] First, for the target as a whole, the target deformation perception Λ reflects the change trend of the target, and naturally the Gaussian mask parameters of the target should be adjusted according to this trend to achieve adaptive spatial regularization. Considering that the target deformation is often generated from the outside to the inside in space, the distribution of the Gaussian mask in space should be adjusted when the target change trend changes significantly, that is:

[0070]

[0071] σ=std[δlog(Λ+1)]

[0072] Among them, μ X and μ Yrepresents the mean of the two coordinates, σ represents the standard deviation of the normalized deformation perception Λ, and δ is the scaling factor. Here, the normalized standard deviation of the target deformation perception Λ is used as the standard deviation of the basic Gaussian mask. When the target deformation trend increases and outliers are generated in Λ, its standard deviation increases, making the basic Gaussian mask more "wide" in space, thereby increasing the model's learning of outer pixels; conversely, the basic Gaussian mask is more "narrow" in space, thereby reducing the learning of outer pixels, completing the adaptive adjustment of the spatial domain regularization according to the target deformation trend.

[0073] Furthermore, it is not enough to just change the spatial distribution of the overall adaptive Gaussian mask. It is necessary to further refine and modify the local position of the target that produces a large change trend. That is, the mask of the key learning area further considers the deformation trend and spatial distribution of the target on the basis of the adaptive Gaussian, and uses one or more Gaussian distributions to approach the outlier distribution in the key area of ​​Λ, so that the model can learn the key area more finely. The specific steps are as follows.

[0074] First, the target deformation perception Λ is masked. Since large interference is easily generated at the boundary, resulting in noisy outliers in the Λ value, it is necessary to first use a standard Gaussian filter to process Λ to prevent the spatial regularization from weakening the inhibitory effect of the boundary effect in the model.

[0075] Secondly, this paper uses the Z score method to extract outlier samples in Λ. The Z score value of each pixel is calculated as follows:

[0076]

[0077] Among them, i and j represent the pixel coordinate position, μ Λ represents the average value of Λ, σ Λ Represents the standard deviation of Λ. According to conventional statistical experience, the Λ value of pixels with an absolute value of Z score greater than 3 is identified as an outlier sample. Considering that the deformation area is often connected in space, the obtained discrete samples are spatially clustered using the DBSCAN method. By measuring the distribution of data in each category obtained by clustering, including extreme values, means and standard deviations in two coordinate directions, a two-dimensional Gaussian distribution model of each category is constructed. Here, the two coordinates are considered to be independent of each other, that is, the correlation coefficient in the two-dimensional Gauss is 0. The formula is as follows:

[0078] f(X,Y)=max(f1,f2,…,f K )

[0079]

[0080] Where k represents the index of each category. The local refined Gaussian mask f is obtained by taking the maximum value of each two-dimensional Gaussian distribution, which is added and fused with the basic Gaussian mask to obtain the final spatial domain regularization mask W=w+f.

[0081] Adaptive temporal regularization: During the tracking process, the target difference between two adjacent frames determines the difference in model learning results. Since the image sampling frequency in real scenes is high, the model learning can be suppressed when there is no tracking anomaly and the target deformation is not obvious. On the contrary, when the target changes drastically, the model learning should be accelerated. Therefore, different weight parameters can be adaptively assigned to the temporal regularization term according to the target deformation speed, thereby improving the generalization of the model in different situations. Specifically, when the target changes drastically and the deformation perception Λ is large, the penalty coefficient θ in the objective function t The smaller it is, the better it can learn the details of the deformed target.

[0082]

[0083] Where σ is the standard deviation of the normalized deformation perception Λ calculated above, is the time domain regularization coefficient θ t The reference value is calculated as v and v are hyperparameters.

[0084] Step 6. Input the response map and the target deformation trend perception vector into the tracking anomaly detection network based on feature fusion to determine whether the current tracking state is normal.

[0085] The basic idea of ​​this network is to weight the response graph morphology and the response graph change rate and then add and fuse them. After that, it is normalized into probability distribution through convolution, full connection, and SoftMax to complete the classification task, that is, to determine whether there is a tracking anomaly.

[0086] Specifically, the input of the network is a single-channel response map and a single-channel change rate matrix of size H×W. The two inputs are each passed through a convolutional layer with the same output dimension to obtain a C×H×W feature map. After adding and fusion, the two are respectively learned through two branches to obtain the weights after the channel features and global features are fused. Among them, the channel feature branch passes through the pixel convolution layer to obtain the deep features of C1×H×W, and then passes through the ReLU activation function and another pixel convolution layer to restore the feature dimensions to C×H×W, so as to align the feature dimensions during weighted fusion. After the global feature branch passes through the global average pooling, the pixel convolution layer obtains the deep features of C1×H×W, and passes through the ReLU activation function and another pixel convolution layer to restore the feature dimensions to C×H×W. The global feature branch makes the network focus more on global information, while reducing the amount of data and reducing the risk of overfitting. The feature maps output by the two branches are added and fused, and then mapped to between 0 and 1 through the Sigmoid activation function. This is used as the weight of the response map distribution, and the corresponding weight of the response map change rate is naturally obtained. After the two are weighted and fused according to the weights, further deep feature extraction is performed through the convolution layer, and then the output of the activation function is converted to a probability distribution through the fully connected layer through Softmax. Finally, the classification of the current response map is obtained based on the probability of each category, and the judgment of whether there is an abnormality in the tracking process is completed.

[0087] It is worth noting that under different scenes, different targets, different lighting conditions, etc., tracking anomalies such as target occlusion have very similar morphological distributions on the response graph, which makes the anomaly detection network that integrates the response graph distribution and the target change rate itself have good robustness and generalization capabilities. Therefore, compared with training directly on the original image, training on the response graph can achieve better classification results and generalization capabilities while using less training data and a simpler network structure. In this model, the tracking anomaly detection network is only trained and tested on a 1485-frame video in the VisDrone dataset at a ratio of 7:3. The resulting network brings comprehensive performance improvements in the entire dataset test.

[0088] Step 7. Filter Update

[0089] If there is no tracking anomaly, the spatial mask and temporal regularization coefficients obtained in step 5 are updated to the objective function, a new filter is calculated by the ADMM algorithm, and the original filter in step 2 is updated; if there is a tracking anomaly, the original filter in step 2 is kept unchanged. After completing the tracking or keeping of the filter, wait for the next frame of image input.

Claims

1. A UAV target tracking method based on target deformation and feature fusion, characterized in that: The following steps are involved: 1) Obtain the target image through the drone’s onboard camera; 2) Construct a tracking algorithm model and use the model to calculate the position of the target to be tracked in the image; 3) Update the filter in the tracking algorithm model and wait for the next frame of image input.

2. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 1, characterized in that: The step 2) comprises the following steps: 2.1) The current frame image and the filter are transformed into the frequency domain by the fast Fourier algorithm and then filtered to obtain a response graph; 2.2) The peak position of the response graph is used as the position of the target to be tracked in the current frame.

3. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 1, characterized in that: The step 3) comprises the following steps: 3.1) The response map is stored in the response map vector, and the target deformation trend perception vector is calculated; 3.2) Calculate the target spatial domain mask and temporal domain regularization term coefficient according to the target deformation trend perception vector; 3.3) Input the response graph and the target deformation trend perception vector into the tracking anomaly detection network based on feature fusion, and judge whether the current tracking state is normal according to the detection results; 3.4) If there is no tracking anomaly, the spatial domain mask and the temporal domain regularization term coefficients are updated to the objective function, a new filter is calculated by the ADMM algorithm, and the original filter is updated; if there is a tracking anomaly, the original filter is kept unchanged.

4. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 3, characterized in that: The calculation target deformation trend perception vector is specifically: The second-order response diagram is taken as the rate of change vector of the response diagram Λ=[|Λ 1 |,|Λ 2 |,|Λ 3 |,…,|Λ T |], the i-th element |Λ i |For: in, represents the i-th element of the t-th frame response graph, Φ Δ represents an alignment shift operation, represents the change of the i-th element of the response graph of the t-th frame, Λ i Indicates the change rate corresponding to the current response graph, that is, the target deformation trend perception vector.

5. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 3, characterized in that: The target spatial domain mask W is specifically: W=w+f Among them, w is the basic Gaussian mask: σ=std[δlog(Λ+1)] Among them, μ X and μ Y represents the mean of the two coordinates, σ represents the standard deviation of the normalized deformation perception Λ, δ is the scaling factor, std represents the standard deviation calculation function, X, Y represent the horizontal and vertical coordinates; f is the local refined Gaussian mask: f(X,Y)=max(f1,f2,…,f K ) Among them, k represents the index of each category, and each two-dimensional Gaussian distribution is optimally obtained to obtain a local refined Gaussian mask f.

6. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 3, characterized in that: The time domain regularization term coefficient θ t Specifically: in, is the time domain regularization coefficient θ t , ReLU represents the ReLU function, and ζ and v are hyperparameters.

7. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 3, characterized in that: The step 3.3) comprises the following steps: 3.3.1) The single-channel response map and the single-channel change rate matrix of size H×W are respectively passed through a convolutional layer with the same output dimension to obtain a C×H×W feature map; 3.3.2) Add and fuse the two feature maps; 3.3.3) Learning channel features and global features through channel feature branches and global feature branches respectively; 3.3.4) The feature maps output by the two branches are added and fused, and then mapped to the range of 0 to 1 through the Sigmoid activation function, which is used as the weight of the response map distribution; 3.3.5) After the single channel response map and the single channel change rate matrix are weighted and added according to the weights, the deep feature extraction is further performed through the convolution layer; 3.3.6) The weighted features are used as the input of the activation function, the feature values ​​are output, and the Softmax is converted into a probability distribution through the fully connected layer; 3.3.7) The classification of the current response graph is obtained according to the probability of each category, and the judgment of whether an abnormality occurs during the tracking process is completed.

8. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 7, characterized in that: The channel feature branch performs the following steps: The channel features are passed through the pixel convolution layer to obtain the deep features of C1×H×W, and then after passing through the ReLU activation function, another pixel convolution layer restores the feature dimension to C×H×W.

9. The method for tracking unmanned aerial vehicle targets based on target deformation and feature fusion according to claim 7, characterized in that: The global feature branch performs the following steps: After the global features are pooled through global average, the deep features of C1×H×W are obtained by the pixel convolution layer. After the ReLU activation function, another pixel convolution layer restores the feature dimension to C×H×W.