Precise Tracking Method for Moving Objects in Consumer Drone Videos
By adopting decision-aware correlation lock tracking algorithm, difference superposition detection method and two-dimensional scale estimation technology based on drone video target tracking, tracking difficulties caused by target deformation, occlusion and scale changes are solved, and a higher tracking accuracy and success rate are achieved.
Patent Information
- Application Number
- CN202210296764.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-25
AI Technical Summary
In the face of factors such as target deformation, occlusion, and scale changes in consumer-grade drone videos, it is difficult to achieve robust and reliable target tracking, resulting in tracking drift and failure.
An association locking tracking algorithm based on decision perception is proposed, combining feature association selection, feature decision perception and decision perception weight calculation, and multi-feature description targets are used to reduce tracking drift caused by deformation; at the same time, an association locking tracking method based on differential superposition detection is proposed, and the tracking effect under occlusion is improved through fast locking of moving targets and differential superposition detection; a two-dimensional target robust scale estimation method is proposed, and the scale changes of the target are accurately estimated through perspective projection model and the two-dimensional scale estimation tracking framework of correlation locking is estimated.
It significantly improves the accuracy and success rate of drone video target tracking, can better cope with complex scenarios such as target deformation, occlusion and scale changes, and improves the overall tracking accuracy and performance.
Smart Images

Figure CN114627156B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for tracking moving objects in drone videos, and particularly to a method for accurately tracking moving objects in consumer drone videos, belonging to the technical field of drone video object tracking. Background Art
[0002] Object tracking is to analyze a series of consecutive video images obtained by visual sensors such as cameras, and obtain information such as the position, size, and motion state of a specific object or multiple objects. Visual tracking is to predict the subsequent state of the object by using the initialized target box in the image sequence. Compared with common computer vision problems, visual tracking is special because of very limited positive samples and infinite negative samples.
[0003] The proposal and application of intelligent transportation and driverless vehicles, as well as the rapid development of drones, have greatly promoted the development and application of object tracking algorithms. During the tracking process, due to some external factors, such as changes in illumination, changes in shooting perspectives, object occlusion, etc., the main challenges in current tracking are: deformation of the object caused by its own motion or external factors, occlusion by other objects during motion, changes in scene illumination, and motion blur caused by object motion. Due to the variability of objects and scenes in the real world, there is no reliable method in the existing technology to solve object tracking in any scenario. Therefore, object tracking for drone video surveillance platforms, which is one of the current research and development and application hotspots, on the one hand, currently the drone technology is developing rapidly and the application of drones is becoming more and more extensive, gradually penetrating into various scenarios of production and life. Object tracking based on this platform has an urgent market demand; on the other hand, when the drone is moving, the object is prone to deformation and undergoes a drastic scale change in a short time, which is also a difficulty in object tracking, making the tracking algorithm in this application background applicable to other application scenarios.
[0004] The continuous improvement of computer performance in recent years has provided a basic platform support for the development of the visual tracking field. The continuous emergence of new methods and the cross-reference of methods from other disciplines have promoted the continuous development of visual tracking, and many research and development results have emerged.
[0005] Target appearance model of the prior art: It includes a target appearance description and a model learning method. Among them, the appearance description solves the problem of how to design a descriptor with strong robustness and distinctiveness, while the model learning method solves the problem of how to learn the apparent model of the target from existing samples. In image processing, features are extracted to describe an image. The features include color features, intensity information, texture characteristics, contour information, edge gradients, etc. Some more complex ones are scale-invariant features. Features are extracted from local regions of the image to reflect changes in attributes such as image intensity and color. Some of these features can overcome adverse effects such as changes in illumination and rotational poses of the target. Under the framework of correlation filters, fusing multiple convolutional layers is still an open problem. Tracking based on fusing multiple convolutional layers using correlation filters uses too many parameters and is prone to overfitting.
[0006] For the design of the target appearance model, one approach is to only focus on the target itself, and the other is to model and analyze both the target and the background. Since the appearance of the tracking object may vary greatly in a video, it is not advisable and unreliable to estimate the model only based on the first frame image and use this single model to locate the target in the remaining images. Generally speaking, good tracking algorithms will use useful information in the image to adaptively update the model. For example, the tracking results of each frame are used as training data to update the model. In online tracking, the current tracking result is often used as a positive sample (sometimes extended to very nearby positions as positive samples), and the surrounding area of this position is used as a negative sample, and then an adaptive appearance model is built. Generally speaking, there are two main categories of model learning methods: generative models and discriminative models.
[0007] Target motion model of the prior art: It is used to estimate the possible positions where the target may appear. It is based on the continuous movement of the object in the field of view and no sudden change in speed. According to the position information of the target in the current frame and previous image frames, the target motion trajectory is obtained, so as to estimate the possible position information of the target in the next frame. Common algorithms include: optical flow method, Kalman locking method, particle locking, median drift method.
[0008] Update of the target model in the prior art: By modeling the target appearance, the target is distinguished from the background. Through the motion estimation of the target, the position of the target in the next frame of the image is obtained. The model update is closely related to whether the algorithm can handle problems such as target deformation and occlusion. A relatively direct method is template replacement. Using the idea of matching, after obtaining the target in the current image, the tracking result is used as an observation sample to update the existing model. When using this method, it is necessary to evaluate the reliability of the template. Otherwise, once the target is occluded, the tracking will drift. The learning method based on the incremental subspace constructs the target model with the latest obtained tracking result. The template update methods based on template replacement and incremental subspace learning are both based on the generative model, while the discriminative model regards the target model update as a classification problem, and the core is to design a classifier. These model update methods all face the same problem: it is difficult to determine the reliability of the samples. Unreliable training samples will calculate background noise into the target appearance model, resulting in contamination and model degradation.
[0009] In summary, there are still problems in the prior art of UAV video target tracking. The difficulties and problems to be solved in this application mainly focus on the following aspects:
[0010] First, during the target tracking process of consumer UAVs, there are some external factors that increase the tracking difficulty, including light changes, target deformation, motion blur, occlusion, etc. Due to the variability of objects and scenes in the real world, there is no reliable method in the prior art to solve target tracking in any scenario. The difficult problems in the application scenarios of robust tracking on the UAV motion platform are prominent. Compared with the video obtained by the traditional pan-tilt, in the video obtained by the consumer UAV, the target is more likely to have large deformations and scale changes in a short time, which undoubtedly increases the tracking difficulty. Single features have limitations for complex environments, and the tracking effect is not ideal. The limited prior information about the target during the tracking process, and the tracking model must be robust to the appearance changes of the object. The former is likely to make the model too complex, resulting in too large a test error while pursuing a small training error. The latter is because the target is not constant during the tracking process, and factors such as light, occlusion, motion blur, and rotation all cause changes in the target. For the target deformation during the tracking process, the prior art lacks effective solutions.
[0011] Second, the existing technology uses a circulant matrix to achieve dense sampling. However, in the tracking environment of consumer drones, the target may deform or even experience motion blur, resulting in errors in tracking after constructing the circulant matrix, thus leading to drift. The uncertainty and complexity of the consumer drone scenario pose many challenges to target tracking. These problems that may be encountered during the tracking process are not conducive to forming a robust and reliable visual tracking. Among them, the occlusion problem is a difficult problem in drone target tracking. First, the reasons for occlusion are complex and diverse, without prior knowledge. The occlusion time and degree are unpredictable. Partial occlusion and full occlusion have different effects on the target appearance model. Full occlusion will seriously contaminate the target appearance. Once the target is occluded, it will directly affect the update of the target appearance and motion estimation model. Second, it is difficult to identify occlusion itself. It is easy for the human eye to judge that the target is occluded, but this is a major problem for machine learning. If the occluded area cannot be identified, once the occluded samples are used in sample update, the target appearance model will be affected and degraded or even contaminated, which may lead to tracking drift. How to reduce the impact of occlusion on sample update and even the entire tracking process is an urgent problem to be solved.
[0012] Third, the existing technology using a circulant matrix has well solved the problem of large computational complexity. Using the fast Fourier transform, the training and detection computational complexity is significantly reduced, which is especially suitable for tracking. However, there are also some disadvantages: First, the input image is circularly shifted, and a circular sliding window is used to train the classifier on the training set. However, due to this periodic assumption, there will be periodic overlap in the training images. Due to the discontinuity of the sample edges, obvious segmentation lines will be generated after circular shifting. These segmentation lines appear as high-frequency noise in the frequency domain, which has an adverse effect on frequency domain processing. Second, the target response in the filter training stage is independent of the observed image. When in scenarios such as fast motion, occlusion, and motion blur, the circular motion does not correspond one-to-one with the actual translation of the target. The samples used to train the tracking filter are not manually labeled but are based on the tracking results of the previous frame. Due to the markings of the algorithm itself, these factors inadvertently use the already biased images as positive samples for learning and training. The target detection part of the tracker is inaccurate, which has an adverse effect on the target position information of the next frame image and affects the parameter update of the filter. Finally, there is a potential problem of tracking drift. When the target rotates or deforms, it is easy to have inaccurate tracking results. When the target is semi-occluded or even fully occluded, the target appearance model loses its function, directly leading to tracking failure. In the case of motion blur causing target distortion, it seriously damages the training samples, making the target appearance model lose its discrimination ability and reducing the tracking accuracy.
[0013] Fourth, the strategy adopted by the prior art to determine the target position information is that the position corresponding to the maximum filter response value is the location of the target center. However, this sample update method has the following two disadvantages: First, regardless of whether the current image can be accurately tracked, its tracking result will be accumulated to the next frame of the image, affecting the subsequent tracking effect; the circular motion does not correspond one-to-one with the actual translation of the target. The circular motion is only an approximation of the actual translation in the image. In real tracking scenarios, especially with fast movement and occlusion of express delivery, it becomes unreliable. Using a single center Gaussian as the target response hinders the performance of the tracker and causes irrecoverable drift. Second, the prior art tracking algorithm uses continuous and deterministic weight sample updates in sample updates, that is, the results obtained in each frame will affect the subsequent target, and the sample update weights will be directly determined in the initialization. Then, the influence of each frame of the image on the tracking result has also been determined. However, in the actual tracking scenario, the role of each frame of the image in sample training is different. Especially when the target is occluded, the target appearance features extracted at this time are quite different from the original target, basically the target background. Once misidentified as the target and added to the sample training, it will contaminate the training samples. Summary of the Invention
[0014] This application takes the visual tracking of an unmanned aerial vehicle (UAV) motion platform as the scenario and makes improvements in three aspects of deformation, occlusion, and scale based on the correlation-locked tracker: First, a correlation-locked tracking algorithm based on decision perception is proposed. Based on the tracking framework of decision perception, through feature correlation selection, feature decision perception, and decision perception weight calculation, fusion is carried out from coarse to fine at the decision layer, and the target is described with multiple features to reduce the tracking drift caused by the deformation of the target due to various factors. Second, to solve the tracking drift and even tracking failure after the target is occluded, a correlation-locked tracking method based on difference superposition detection is proposed, through rapid locking of the moving target and correlation-locked target tracking based on difference superposition detection. Third, a two-dimensional target robust scale estimation method is proposed. The tracking effect of this application has been significantly improved in terms of both tracking accuracy and success rate. Using two-dimensional scale estimation can more accurately estimate the scale of the target, thereby improving the overall tracking accuracy and enabling the algorithm to have better tracking performance.
[0015] To achieve the above technical feature advantages, the technical solutions adopted by this application are as follows:
[0016] A precise tracking method for moving targets in consumer drone videos, which is improved in terms of deformation, occlusion, and scale based on an association-locking tracker: First, an association-locking tracking algorithm based on decision perception is proposed, including: one is a tracking framework based on decision perception; the second is feature association selection; the third is feature decision perception; the fourth is decision perception weight calculation; Second, an association-locking tracking method based on difference superposition detection is proposed, including: one is rapid locking of moving targets; the second is association-locking target tracking based on difference superposition detection; Third, a two-dimensional target robust scale estimation method is proposed, including: one is scale estimation based on a perspective projection model; the second is a two-dimensional scale estimation tracking framework based on association locking; the third is a two-dimensional scale prediction and evaluation strategy; the fourth is adaptive search window scale selection; the fifth is the process of scale prediction and estimation method;
[0017] (1) For the drift caused by target deformation during tracking, different features are used for tracking respectively and fused at the decision layer to obtain the final result. One method is to construct two association filters, which respectively use the HOG feature with target structure information and the CNs with target color information. After calculating the tracking results, the maximum response value of the filter is used at the decision layer to determine the weight of each filter for the final result. Another method is to use an association-locking and a tracking framework based on color histogram, and then fuse the results of the two to obtain the final tracking result;
[0018] (2) For the target occlusion problem during tracking, an association-locking tracking method based on difference superposition detection is proposed. First, the structural risk function in ridge regression is replaced with a regularization function, and components are added to better handle the occlusion problem. Then, the difference superposition features are extracted within the search box and superimposed with the original samples to increase the difference between the target and the surrounding background, and at the same time, the potential risk of the circulant matrix is excluded. A method based on the detection result is adopted for sample update, and non-continuous update is used for model update to handle the target occlusion problem during tracking;
[0019] (3) For the target scale change problem during tracking, a two-dimensional target robust scale estimation method is proposed. First, the rough scale estimation of the target is obtained by using sensor data and the internal parameters of the shooting camera. A two-dimensional association filter is designed to estimate the target scale, and the length and width dimensions of the target are estimated separately. When the target deforms and scales change, the true size of the target in the field of view can be accurately estimated. In addition, an adaptive search window method is proposed to change the size of the search box according to the displacement of the target in adjacent two frames to prevent the target from moving too fast and moving outside the search window; the samples are downsampled on the scale target filter to reduce the amount of calculation data.
[0020] Preferably, a tracking framework based on decision perception: achieving the optimal tracking effect with a model as simple as possible and a small number of parameters. Suppose in the t-th image I t , the tracking problem is regarded as finding the most likely (highest-scoring) region p from the set of candidate regions S t : t
[0021]
[0022] where T refers to performing a certain transformation on the image, and f(T(I t , p); θ) represents the score of the rectangle p in the image I t under the model parameters θ. The model parameters satisfy minimizing the loss function:
[0023]
[0024] where the loss function L(θ; χ t ) is related to the previous image and the position of the target in the image. is the entire model parameter space, R(θ) and λ are regularization factors to prevent the model from being too complex and causing overfitting. Achieving effective tracking is transformed into the problem of selecting the functions f and L;
[0025] When the target is initialized, the position and scale information of the target are obtained. Finding the region most similar to the target in the candidate image sequence is the tracking target. Judging whether two signals are correlated is to measure the similarity between them. When calculating the similarity between the candidate region sample and the target, the two are subjected to a correlation operation. The loss function is selected to calculate the distance between two vectors. The tracking algorithm uses positive samples containing the target and negative samples that are all backgrounds, and uses ridge regression to construct a regressor to train the appearance model of the target.
[0026] Preferably, feature correlation selection: Based on complementary features, multi-feature decision perception is performed on the moving target, and according to the different impacts of each feature on the tracking result, a framework is designed for fusion to obtain a more accurate and stable tracking effect;
[0027] Among them, for HOG feature selection, the sample region is divided into several sub-regions, and then 32-dimensional features are extracted from each region. Excluding the last dimension which is all 0 and not considered, the remaining 31-dimensional features are used. For texture features, the RFS filter bank of the fast anisotropic Gaussian filter is used, and this feature has a total of 8 dimensions.
[0028] Preferably, for feature decision perception: Two filters are trained using different input samples. One of the input images is the target box and a surrounding circle of background, which contains spatial context information, i.e., with a size of window_sz. To better distinguish the target and background by differential superposition, HOG features are selected. In addition, to eliminate the boundary effect, a cosine window function is applied to the extracted features. The other input image only contains the target and is smaller than the size of the target itself, i.e., with a size of sz. To enable the tracker to still have good robustness when the target undergoes large deformations, CNs features are used. The former incorporates background information to expand the search area for the target, while the latter focuses only on the target to improve accuracy. The two filters update the template online, making the tracking result more accurate. By using multiple feature decision perceptions for the target, more target information is retained and the target is tracked through multiple means, better coping with unexpected situations during the tracking process.
[0029] Preferably, for decision perception weight calculation: A multi-feature perception additive model is adopted. Under the additive fusion strategy, the final results after n decision perceptions are additively fused. Additive fusion is insensitive to noise and increases the robustness of the tracking algorithm. According to the filter response value and the target position, the final target position is obtained as:
[0030]
[0031] where res1 and res2 are the maximum response values of the two filters, and p 1 and p 2 are the positions of the target corresponding to the maximum response values respectively;
[0032] The algorithm flow for HOG and CNs decision perception:
[0033] Input: Image sequence Initial position p of the target 0 ;
[0034] Output: Position information of the target in each frame of the image
[0035] Process 1: Initialization, t = 1;
[0036] Process 2: Use I 0 and p 0 to initialize trackers Track1 and Track2;
[0037] Process 3: for t = 1:T
[0038] Process 4: Calculate the current search sub-window according to the output p t-1 from the previous frame;
[0039] Process Five: For tracker Track1, HOG features are adopted. Using the target structure information, the search window contains spatial context information, and a Hanning window is used for window processing.
[0040] Process Six: For tracker Track2, CNs features are adopted. Using the global color information of the target, the search window is the target itself and no window processing is performed.
[0041] Process Seven: Response maps of the two trackers are obtained from the tracking model.
[0042] Process Eight: The maximum response value is found in the response map, and this value corresponds to the most likely location of the target.
[0043] Process Nine: Fusion is performed by the additive model of multi-feature perception to obtain the final tracking result p of the target. 0 ;
[0044] Output: The position information of the target in each frame of the image.
[0045] Preferably, for fast locking of a moving target: The tracking algorithm for fast locking of a moving target uses a filter W to associate with an alternative region to estimate the target position. Among them, an image block that contains the target and is larger than the target is used as the search window, and then the search window is circularly shifted to obtain different alternative windows. The actual displacement of the target is approximated by circular motion, and finally a data matrix with a circular structure is constructed to obtain training samples.
[0046] Assume the kernel matrix K z is obtained from the kernel correlation operation of the matching template and the search window samples, where the matching template X and the detection sample Z are circulant matrices obtained by shifting the vectors x and z. K z is also a circulant matrix:
[0047] K z = C(k xz ) Equation 4
[0048] where k xz is the kernel correlation operation of the vectors x and z. After introducing the kernel function in the least squares, the regression function is f(z):
[0049] f(z) = (K z )α Equation 5
[0050] α is the transformation coefficient. Performing DFT operations on both sides gives:
[0051]
[0052] Among them, z is the search window image patch that predicts including the target, from which the position of the target is detected, x is the target model learned in the image, the right side of the equation is a linear combination of k multiplied by the transformation coefficient α, ⊙ is the element-by-element multiplication in the vector, * is the conjugate, perform the inverse DFT operation on Equation 6 to obtain the response matrix of the filter, and the position corresponding to the maximum value in the matrix is the target position. The model is updated as follows:
[0053]
[0054] Among them, η is the learning factor, and are the Fourier transforms of the coefficient α of the current frame t and the previous frame t - 1 respectively, and represent the Fourier transforms of the target matching templates of the current frame and the previous frame respectively.
[0055] Preferably, correlation-locked target tracking based on differential superposition detection: further improve in two parts: sample training of the filter and filter parameter update;
[0056] 1. Differential superposition detection
[0057] Add a regularization parameter to prevent model overfitting and reduce the test error. Use sparse parameters to effectively handle occlusion. In the training sample stage, take the differential superposition area into consideration, and give priority to processing the parts that are easy to attract attention, which is convenient to improve the tracking performance when the target is occluded. Focus on the background, remove the statistical redundancy of the input signal, remove the redundant information of the target to obtain the significant target of the image, and use the log spectrum of the image. The log amplitude spectrum of an image minus the average log amplitude spectrum is the differential superposition part of the image.
[0058] While considering the differential superposition information of the image, retain the original information of the target. During the sample training process, Ω(u) = u + S BB (t), where u is the input image, and S BB (t) is the significant area;
[0059] After removing the redundant information of the image, perform deeper differential superposition detection, which is represented by a Boolean mapping based on the instantaneous awareness of the scene. According to the prior distribution and random threshold of the feature space, use the feature map of the input image to generate a series of Boolean maps:
[0060] B i = THRESH(φ(I), θ),
[0061] φ ~ p φ , θ ~ p θ Equation 8
[0062] Among them, the function THRESH(*,θ) binarizes the image through θ, φ(I) represents the feature map of the image I, and is normalized to the range of 0 - 255, p φ and p θ represents the prior distribution, and the influence of the boolean map on visual attention is represented by the attention map A(B):
[0063]
[0064] where I is the input image, is the average attention map, which is used for subsequent processing to obtain a final difference superposition map S;
[0065] 2. Target sample update
[0066] During the training phase, the response value is considered in the weight of sample update, and the results detected by the tracking algorithm are utilized, rather than simply adopting a binary strategy. First, the peak sidelobe ratio is used to determine whether the tracking is effective, and then the influence of the current tracking result on the next frame tracking process is directly determined according to the maximum response value corresponding to the tracking result. Assuming that starting from the second frame image, the maximum response value of each image is stored in response_all, the update of the sample is as follows:
[0067]
[0068] where η is the learning factor, and respectively represent the Fourier transforms of the target matching templates of the current frame and the previous frame. The detection results are considered in the sample update. When the target is contaminated, the occupied weight decreases, and the influence on the subsequent target also decreases;
[0069] 3. Tracking model update
[0070] When updating the model, a sparse update strategy is adopted. The model is updated every Ns frames, and the samples still need to be updated every frame. The model update frequency is reduced, saving time and avoiding the problem of model drift. Finally, in this application, Ns is taken as 6.
[0071] Preferably, a two-dimensional scale estimation tracking framework based on correlation locking: Design two consistent correlation filters, defined as the position target filter and the scale target filter, which respectively implement target tracking and scale transformation. The former is used for the positioning of the target in the current frame, and the latter is used for the scale estimation of the target in the current frame. The two filters are relatively independent, and different feature types and feature calculation methods are selected for training and testing. A two-dimensional filter is adopted, assuming that the length and width of the target do not change in the same way. Even if the target undergoes large deformation, the target scale can still be accurately estimated.
[0072] The overall framework of the tracker incorporating two-dimensional target robust scale estimation is divided into two parts: position prediction and scale estimation. The specific steps are as follows: Training samples are obtained through dense sampling in the search area around the target and feature extraction is performed. The samples are subjected to Fourier transform, and a filter is trained using least squares regression in the frequency domain; An online learning method is used to obtain a position kernel correlation filter. Then, the output value of the filter is mapped to the time domain through inverse discrete Fourier transform, and the coordinates of the maximum point in the response map are found. This point corresponds to the center position of the target in the sample; Then, the scale prediction process based on the correlation filter is similar to the position process of the predicted target. After obtaining the position of the target, according to the preset scale value, downsampling and upsampling are performed around the target to obtain a series of image patches of different scales; These image patches are then bilinearly interpolated to change the size of the image patches to be the same as the designed scale model to obtain training samples; Next, feature extraction is performed. After obtaining the HOG features of the image patches, a least squares classifier is trained to obtain a scale target tracker; In addition, after obtaining the features of the image patches, the features are windowed using a Hamming window to suppress the high-frequency noise caused by the image boundary in the frequency domain due to the use of a circulant matrix;
[0073] When applying the scale target tracker to a new image, the score in the scale space is calculated, that is, the response value of the filter. The position of the maximum value in the scale target filter response map corresponds to the final scale of the target;
[0074] A position target filter is constructed to obtain the position of the target to be tracked. A scale target filter is constructed and the target scale is obtained using the center position of the target obtained by the position target filter. The final target information is passed to the position target filter of the next frame of the image, and the two operate together to update the parameters online and improve the performance of the entire tracker.
[0075] Preferably, the two-dimensional scale prediction evaluation strategy: A two-dimensional scale correlation filter is learned to detect the scale change of the target, and the target search area is reasonably limited according to the scale change of the target;
[0076] The most important part of target scale estimation is to give specific alternative scale transformation values for the target: In the scale prediction evaluation process, for the current image, the size of the target is P×R, and the size of the scale target filter is S×S. For each Image patches J of size are extracted separately in the two dimensions of the length and width of the target where a is the filter parameter factor. For targets with large scale changes, the alternative scales should be as different as possible. Conversely, the alternative scales should change less. n×n
[0077] Adaptive search window scale selection: The size of the search box is constrained according to the offset of the target positions in two adjacent frames. Based on the tracker of the association locking framework, when performing the association operation, the scale of the filter is initially determined during initialization. The size of the search area is changed according to the offset of the target positions in two adjacent frames, and the size of the search window is adaptively selected by comparing the offsets of the target centers in two adjacent frames. A critical value is preset. When the offset is greater than this critical value, the search window is enlarged.
[0078] Preferably, the process flow of the scale prediction and estimation method: In the entire tracking algorithm after adding scale estimation, first obtain the position information using the position target filter, and then obtain the current scale change of the target using the motion scale target filter. The kernel function involved in the algorithm uses the Gaussian kernel function;
[0079] In the scale estimation module, HOG features are used. First, the features of the image are extracted to obtain the training samples f of the d-dimensional features of the image l , where l = {l, 2, …, d}, and each feature dimension corresponds to an association filter h l , minimizing the loss function:
[0080]
[0081] where g is the output of the filter designed for the training sample f, and λ ≥ 0 is the regularization parameter to control the structural error. Solving the equation gives the filter as:
[0082]
[0083] Adding a regularization factor, performing an l 2 constraint on the filter, so that f has no non-zero terms, alleviating the problem of zeros appearing in the frequency domain of the filter and avoiding the phenomenon of the denominator being zero. Online solve the update problem of a d×d-dimensional linear equation system to obtain a robust approximation and update the association filter H l t for the numerator A l t and the denominator B t :
[0084]
[0085] where η is the learning rate, and the final output of the filter, that is, the association score:
[0086]
[0087] The scale transformation corresponding to the maximum value in the correlation response map is the current scale size of the target;
[0088] When extracting features from the target alternative regions, the length and width of the target are respectively scaled to obtain alternative regions. Although the scale change of the target can be accurately estimated, the computational complexity increases. The principal component feature transformation is used for dimensionality reduction. The eigenvalues and eigenvectors of the covariance matrix of the high-dimensional data are calculated. The eigenvectors are the selected basis vectors, and the eigenvalues corresponding to the eigenvectors represent the projection of the original data on the vectors. The larger the eigenvalue, the more information of the original data is retained by the corresponding eigenvector. The redundant information of the data is removed, and the amount of data to be calculated is reduced;
[0089] In addition, on the scale target filter, the size of the target is concerned. To reduce the computational complexity and retain the overall information of the target, the target is downsampled to a certain size before calculating the target features, rather than processing the entire target.
[0090] Compared with the prior art, the above technical solution has the following innovation points and advantages:
[0091] First, based on the correlation filter, superior performance is shown in the tracking speed. However, for the main challenges in the tracking process, such as the drastic appearance changes caused by deformation, occlusion, scale change, scene illumination change, etc., there are still quite a lot of difficulties. This application takes the visual tracking of an unmanned aerial vehicle (UAV) motion platform as the scenario and makes improvements in three aspects of deformation, occlusion, and scale on the basis of the correlation lock tracker: First, a correlation lock tracking algorithm based on decision perception is proposed. Based on the tracking framework of decision perception, through feature correlation selection, feature decision perception, and decision perception weight calculation, fusion is carried out from coarse to fine at the decision layer, and the target is described by multiple features to reduce the tracking drift caused by the deformation of the target due to various factors; Second, to solve the tracking drift or even tracking failure after the target is occluded, a correlation lock tracking method based on difference superposition detection is proposed, through rapid locking of the moving target and correlation lock target tracking based on difference superposition detection; Third, a two-dimensional target robust scale estimation method is proposed. The tracking effect of this application, whether it is the tracking accuracy or the success rate, has been significantly improved. The two-dimensional scale estimation can more accurately estimate the scale of the target, thereby improving the overall tracking accuracy and making the algorithm have better tracking performance;
[0092] Second, for the drift caused by the deformation of the target during tracking, the decision-making perception of this application combines multiple features to describe the target from multiple aspects. Finally, different features are used for tracking respectively, and fusion is performed at the decision-making level to obtain the final result. One of them is to construct two correlation filters, using the HOG feature with target structure information and the CNs with target color information respectively. After calculating the tracking results, the maximum response value of the filter is used at the decision-making level to determine the weight of each filter for the final result. Another method is to use correlation locking and a tracking framework based on color histograms, and then fuse the results of the two to obtain the final tracking result. In the case of target deformation, fusing at the decision-making level has better tracking performance than fusing at the feature level. Using correlation locking with multi-feature perception for target tracking solves a series of problems brought by large deformations and scale changes of the target in a short time in consumer drone videos, which has great significance and huge practical value.
[0093] Third, for the problem of target occlusion during tracking, this application improves target tracking in the case of occlusion from aspects of target motion estimation and sample update, and proposes a kernel correlation filter tracking algorithm based on difference superposition. First, the structural risk function in ridge regression is replaced with a regularization function, and the l 1 component is added to better handle the occlusion problem. Then, the difference superposition feature is extracted within the search box and superimposed with the original sample to increase the difference between the target and the surrounding background. At the same time, the potential risk of the circulant matrix is eliminated. In terms of sample update, a method based on the detection result is adopted, and drift is reduced in model update, and a discontinuous update method is used to handle the target occlusion problem during tracking. This application also further improves both the sample training of the filter and the update of filter parameters. Currently, the tracking algorithm based on DCF sacrifices tracking speed and real-time performance while pursuing tracking effects. This application improves the tracking algorithm in this link of model update to improve the speed.
[0094] Fourth, for the problem of target scale change during tracking, this application proposes a two-dimensional target robust scale estimation method. First, the rough scale estimation of the target is obtained by using sensor data and the internal parameters of the shooting camera. A two-dimensional correlation filter is designed to estimate the target scale, and the length and width dimensions of the target are estimated separately. When the target deforms and changes in scale, the true size of the target in the field of view can be accurately estimated. In addition, an adaptive search window method is proposed to change the size of the search box according to the displacement of the target in two adjacent frames to prevent the target from moving too fast and moving outside the search window. To improve the calculation speed, downsampling is performed on the samples in the scale target filter to reduce the amount of calculation data. Experimental results in situations such as target deformation, scale change, occlusion, and disappearance that may occur during the tracking process show that this application has better tracking accuracy and overlap rate, and also demonstrates excellent tracking performance when the target is occluded. Brief Description of the Drawings
[0095] Figure 1 It is a perspective imaging schematic diagram for scale estimation based on the perspective projection model.
[0096] Figure 2 It is a central imaging schematic diagram for scale estimation based on the perspective projection model.
[0097] Figure 3 It is a two-dimensional scale estimation tracking framework diagram based on correlation locking.
[0098] Figure 4 It is a comparison diagram of the Groundtruth and different tracking effects of the boat1 test video.
[0099] Figure 5 It is a comparison diagram of the Groundtruth and different tracking effects of the boat2 test video.
[0100] Figure 6 It is a comparison diagram of the Groundtruth and different tracking effects of the boat3 test video.
[0101] Figure 7 It is a comparison diagram of the Groundtruth and different tracking effects of the car4 test video.
[0102] Figure 8 It is a comparison diagram of the Groundtruth and different tracking effects of the wakeboard5 test video. Specific Implementation Method
[0103] In order to make the purpose, features, advantages, and innovation points of this application more obvious, understandable, and convenient for implementation, the following will provide a detailed description of the specific implementation manners in conjunction with the accompanying drawings. Those skilled in the art can make similar promotions without violating the connotation of this application. Therefore, this application is not limited by the specific implementation manners disclosed below.
[0104] The rise and use of consumer drones have brought new application scenarios to target tracking. However, during the tracking process, there are often some external factors that increase the difficulty of tracking, including light changes, target deformation, motion blur, occlusion, etc. This application addresses the difficult problems in the robust tracking application scenario of the consumer drone motion platform and makes improvements in three aspects: deformation, occlusion, and scale based on the correlation lock tracker. First, a correlation lock tracking algorithm based on decision perception is proposed, including: one is a tracking framework based on decision perception, the second is feature correlation selection, the third is feature decision perception, and the fourth is decision perception weight calculation. Fusion is carried out from coarse to fine at the decision level, and the target is described with multiple features to reduce the tracking drift caused by the deformation of the target due to various factors. Second, to solve the tracking drift and even tracking failure after the target is occluded, a correlation lock tracking method based on difference superposition detection is proposed, including: one is the rapid locking of moving targets, the second is the correlation lock target tracking based on difference superposition detection, and the third is to propose a two-dimensional target robust scale estimation method, including: one is the scale estimation based on the perspective projection model, the second is the two-dimensional scale estimation tracking framework based on correlation lock, the third is the two-dimensional scale prediction and evaluation strategy, the fourth is the adaptive search window scale selection, and the fifth is the scale prediction estimation method process;
[0105] (1) To address the drift caused by target deformation during tracking, decision perception is performed at the feature level, multiple features are concatenated, and the target is described from multiple aspects. Finally, different features are used for tracking respectively, and fusion is performed at the decision level to obtain the final result. One method is to construct two correlation filters, using the HOG feature with target structure information and the CNs with target color information respectively. After calculating the tracking results, the maximum response value of the filter is used at the decision level to determine the weight of each filter for the final result. Another method is to use the correlation lock and the tracking framework based on the color histogram, and then fuse the results of the two to obtain the final tracking result; in the case of target deformation, fusion at the decision level has better tracking performance than fusion at the feature level.
[0106] (2) To address the problem of target occlusion during tracking, a correlation lock tracking method based on difference superposition detection is proposed. First, the structural risk function in ridge regression is replaced with a regularization function, and the l 1 component is added to better handle the occlusion problem. Then, the difference superposition features are extracted within the search box and superimposed with the original samples to increase the difference between the target and the surrounding background. At the same time, the potential risk of the subtracted circulant matrix is reduced. A method based on the detection result is adopted for sample update, and a non-continuous update method is used for model update to address the target occlusion problem during the tracking process.
[0107] (3) To address the problem of target scale change during the tracking process, a robust two-dimensional target scale estimation method is proposed. First, a rough scale estimation of the target is obtained using sensor data and the internal parameters of the camera. A two-dimensional correlation filter is designed to estimate the target scale, and the length and width dimensions of the target are estimated separately. When the target deforms and changes scale, the true size of the target in the field of view can be accurately estimated. In addition, an adaptive search window method is proposed to change the size of the search box according to the displacement of the target between two adjacent frames, avoiding the target moving too fast and out of the search window. To improve the calculation speed, downsampling is performed on the samples in the scale target filter to reduce the amount of calculation data. Experiments show that the method of this application has better tracking accuracy and overlap rate, and excellent tracking performance is also demonstrated when the target is occluded.
[0108] I. Decision-Aware Correlation Locked Target Tracking
[0109] Compared with the videos obtained by traditional pan-tilt heads, in the videos obtained by consumer drones, the target is more likely to undergo large deformations and scale changes in a short time, which undoubtedly increases the tracking difficulty. For the target deformation during the tracking process, multi-feature perception correlation locking is used for target tracking.
[0110] (I) Tracking Framework Based on Decision Perception
[0111] The difficulty in the tracking process lies in the limited prior information about the target, and the tracking model must be robust to the appearance changes of the object. The former is likely to make the model too complex, resulting in too large a test error while pursuing a small training error; the latter is because the target does not remain unchanged during the tracking process, and factors such as illumination, occlusion, motion blur, and rotation all cause changes in the target. This application uses a model as simple as possible and a small number of parameters to achieve the optimal tracking effect. Let it be assumed that in the t-th image I t , the tracking problem is regarded as finding the most likely (highest-scoring) region p t from the set of candidate regions S t :
[0112]
[0113] where T refers to performing a certain transformation on the image, and f(T(I t , p); θ) represents the score of the rectangular box p in the image I t under the model parameters θ. The model parameters satisfy minimizing the loss function:
[0114]
[0115] where the loss function L(θ; χ t ) is related to the previous images and the position of the target in the image. For the entire model parameter space, R(θ) and λ are regularization factors to prevent the model from being too complex and causing overfitting, and the problem of effective tracking is transformed into the selection of functions f and L.
[0116] When the target is initialized, the position and scale information of the target are obtained. Finding the region most similar to the target in the alternative image sequence is the tracked target. Judging whether two signals are correlated is to measure the similarity between them. When calculating the similarity between the samples in the alternative region and the target, the two are subjected to a correlation operation. The loss function is selected to calculate the distance between two vectors. The tracking algorithm uses positive samples containing the target and negative samples that are all backgrounds, and a regressor is constructed using ridge regression to train the appearance model of the target.
[0117] (2) Feature correlation selection
[0118] Whether it is the HOG image, color histogram, color attribute or texture feature, it has a certain effect on target tracking, but a single feature has limitations for complex environments and the tracking effect is not ideal. The HOG feature is suitable for describing rigid objects and is insufficient to handle the situation where the target undergoes large deformations; only using color features cannot significantly distinguish the situation where the target and the background have similar colors, and the tracking effect will also be greatly reduced when there is a change in illumination. To address the phenomenon that the target undergoes large deformations during the tracking process, resulting in possible tracking drift, this application is based on complementary features, performs multi-feature decision perception on moving targets, and designs a reasonable framework for fusion according to the different impacts of each feature on the tracking result to obtain a more accurate and stable tracking effect.
[0119] Among them, for the HOG feature selection, the sample region is divided into several sub-regions, and then 32-dimensional features are extracted from each region. Excluding the last dimension which is all 0 and not considered, the remaining 31-dimensional features are used. For the texture feature, the RFS filter bank of the fast anisotropic Gaussian filter is used, and this feature has a total of 8 dimensions.
[0120] (3) Feature decision perception
[0121] Two filters are trained using different input samples. One of the input images is the target box and a surrounding background, which contains spatial context information, i.e., of size window_sz. To better distinguish the target and background by superimposing differences, HOG features are selected. In addition, to eliminate the boundary effect, a cosine window function is applied to the extracted features. The other input only contains the target, and the input image is smaller than the size of the target itself, i.e., of size sz. To make the tracker still have good robustness when the target undergoes large deformations, CNs features are used. The former adds background information to expand the search area of the target, while the latter only targets the target to improve accuracy. The two filters update the template online, making the tracking result more accurate. This strategy uses multiple feature decision perceptions to sense the target, retains more information of the target, and performs target tracking through multiple means to better handle unexpected situations during the tracking process.
[0122] (IV) Decision perception weight calculation
[0123] Using an additive model of multi-feature perception, under the additive fusion strategy, the final results after n kinds of decision perceptions are additively fused. Additive fusion is insensitive to noise and increases the robustness of the tracking algorithm. According to the filter response value and the target position, the final target position is obtained as:
[0124]
[0125] where res1 and res2 are the maximum response values of the two filters, and p 1 and p 2 are the positions of the target corresponding to the maximum response values respectively.
[0126] The algorithm flow of using HOG and CNs decision perception:
[0127] Input: Image sequence Initial position p of the target 0 ;
[0128] Output: Position information of the target in each frame of the image
[0129] Process 1: Initialization, t = 1;
[0130] Process 2: Use I 0 and p 0 to initialize trackers Track1 and Track2;
[0131] Process 3: for t = 1:T
[0132] Process 4: Calculate the current search sub-window according to the output p t-1 of the previous frame;
[0133] Process Five: For tracker Track1, use HOG features, utilize the target structure information, the search window contains spatial context information, and use a Hanning window for window processing;
[0134] Process Six: For tracker Track2, use CNs features, utilize the global color information of the target, the search window is the target itself, and no window processing is performed;
[0135] Process Seven: Obtain the response maps of the two trackers from the tracking model;
[0136] Process Eight: Find the maximum response value in the response map, and this value corresponds to the most likely position of the target;
[0137] Process Nine: Perform fusion by the additive model of multi-feature perception to obtain the final tracking result p of the target 0 ;
[0138] Output: The position information of the target in each frame of the image
[0139] II. Association Locking Tracking Based on Difference Superposition Detection
[0140] The tracking algorithm based on the correlation filter as the basic framework uses a circulant matrix to achieve dense sampling, constructs a circulant data matrix for the cyclic motion of the samples, and based on the properties of the circulant matrix, uses the discrete Fourier transform for fast calculation in the frequency domain, which greatly improves the calculation efficiency. However, in the tracking environment of consumer drones, the target may deform or even have motion blur, resulting in errors in tracking after constructing the circulant matrix, thus leading to drift.
[0141] The uncertainty and complexity of the consumer drone scenario pose many challenges to target tracking, such as scale change, occlusion, appearance change, motion blur, and illumination influence. These problems that may be encountered in the tracking process are not conducive to forming a robust and reliable visual tracking. Among them, the occlusion problem is a difficult problem in drone target tracking. The main reasons are as follows: First, the reasons for occlusion are complex and diverse, without prior knowledge, the occlusion time and degree are unpredictable, and the effects of partial occlusion and full occlusion on the target appearance model are also different. Full occlusion will seriously contaminate the target appearance. This contingency directly affects the update of the target appearance and motion estimation model once the target is occluded. Second, it is difficult to identify occlusion itself. It is easy for the human eye to judge that the target is occluded, but this is a major problem for machine learning. If the occluded area cannot be identified and the occluded samples are used in the sample update, the target appearance model will be affected and degraded or even contaminated, which may lead to tracking drift. Therefore, how to reduce the impact of occlusion on sample update and even the entire tracking process is an urgent problem to be solved.
[0142] The following analysis is based on the problems existing in the update stage of the decision-aware tracking framework. It improves object tracking under occlusion from aspects such as the motion estimation of the target and sample update, and proposes a kernel correlation filter tracking algorithm based on difference superposition.
[0143] (1) Rapid locking of moving objects
[0144] The tracking algorithm based on rapid locking of moving objects uses the filter W to correlate with the alternative regions to estimate the target position. Among them, the image block that contains the target and is larger than the target is used as the search window, and then the search window is cyclically shifted to obtain different alternative windows. The actual displacement of the target is approximated by circular motion. Finally, a data matrix with a circular structure is constructed to obtain the training samples.
[0145] Assume the kernel matrix K z is obtained by performing kernel correlation operation on the matching template and the search window samples. Among them, the matching template X and the detection sample Z are circular matrices obtained by shifting the vectors x and z. K z is also a circular matrix:
[0146] K z = C(k xz ) Equation 4
[0147] where k xz is the kernel correlation operation of the vectors x and z. After introducing the kernel function in the least squares, the regression function is f(z):
[0148] f(z) = (K z )α Equation 5
[0149] α is the transformation coefficient. Performing DFT operation on both sides gives:
[0150]
[0151] where z is the image block of the search window predicting to include the target, and the position of the target is detected from it. x is the target model learned in the image. The right side of the equation is the linear combination of k multiplied by the transformation coefficient α. ⊙ is the element-by-element multiplication in the vector, and * is the conjugate. Performing the inverse DFT operation on Equation 6 gives the response matrix of the filter. The position corresponding to the maximum value in the matrix is the target position. The model update is as follows:
[0152]
[0153] where η is the learning factor, and are the Fourier transforms of the coefficient α of the current frame t and the previous frame t - 1 respectively, and represent the Fourier transforms of the target matching templates of the current frame and the previous frame respectively.
[0154] (II) Advantages and Disadvantages of Circulant Matrices
[0155] The use of circulant matrices solves the problem of large computational complexity. The data is constructed into a circulant matrix, the pixel-by-pixel motion of the target is simulated, and the samples are expanded. According to the properties of the circulant matrix, the amount of information carried by the first row of the matrix can represent the entire matrix. Therefore, there is no need to pay attention to the entire circulant matrix, only the elements of a row or column in it need to be known. In addition, the algorithm uses fast Fourier transform to significantly reduce the amount of training and detection calculations, which is particularly suitable for tracking, where scarce training data and high computational efficiency are crucial for real-time tracking. However, there are also some disadvantages: First, the input image is cyclically shifted, and a cyclic sliding window is used on the training set to train the classifier. However, because of this periodic assumption, the training images will have periodic overlaps. Due to the discontinuity of the sample edges, obvious dividing lines are generated after cyclic shift. This dividing line appears as high-frequency noise in the frequency domain, which has an adverse effect on frequency domain processing; second, the target response in the filter training stage is independent of the observed image. When in scenes such as fast motion, occlusion, and motion blur, the cyclic motion and the actual translation of the target do not correspond one to one.
[0156] The samples used to train the tracking filter are not manually labeled, but the tracking results of the previous frame are used. Based on the labeling of the algorithm itself, these factors inadvertently use the images with existing deviations as positive samples for learning and training. If the target detection part of the tracker is inaccurate (such as fast movement of the target or motion blur), it will have an adverse effect on the target position information of the next frame. Since the target response of the current frame is independent, this error will affect the parameter update of the filter, and there is a potential problem of drift in the final tracking. In this way, when the target rotates or deforms, it is easy to have inaccurate tracking results; when the target is half-occluded or even completely occluded, the appearance model of the target basically loses its function in this case, which directly leads to tracking failure; motion blur may cause target distortion, seriously damage the training samples, and make the appearance model of the target lose its ability to distinguish and discriminate, and the tracking accuracy is reduced.
[0157] Target drift is the main problem in online tracking. The most important reason for drift is that there is a problem with the accuracy of the samples used by the classifier when updating. That is, there is an error accumulation problem in the tracking process. If there is a deviation in the tracking result of the previous frame, the error is accumulated on the classifier, resulting in an incorrect result in the next frame. The visual manifestation is that the target begins to drift and finally the tracking fails.
[0158] 3.3. Correlation-locked target tracking based on difference superposition detection
[0159] This application proposes a tracking algorithm for correlation filters based on differential superposition detection, and further improves it in two parts: sample training of the filter and update of filter parameters.
[0160] 1. Differential superposition detection
[0161] To handle the situation where the corresponding target is occluded, a regularization parameter is added to prevent the model from overfitting and reduce the test error. Sparse parameters are used to effectively cope with occlusion. During the training sample stage, the differential superposition region is taken into account, and the parts that are easy to attract attention are processed first, which is convenient for improving the tracking performance when the target is occluded. The focus is placed on the background to remove the statistical redundancy of the input signal and the redundant information of the target to obtain the salient target of the image. The log spectrum of the image is used. The differential superposition part of the image is obtained by subtracting the average log amplitude spectrum from the log amplitude spectrum of an image.
[0162] To avoid losing the details of the target and retain the original information of the target while considering the differential superposition information of the image, during the sample training process, Ω(u) = u + S BB (t), where u is the input image and S BB (t) is the salient region.
[0163] After removing the redundant information of the image, deeper differential superposition detection is performed. Based on the instantaneous awareness of the scene, there is a Boolean mapping representation. According to the prior distribution of the feature space and the random threshold, a series of Boolean maps are generated using the feature map of the input image:
[0164] B i = THRESH(φ(I), θ),
[0165] φ ~ p φ and θ ~ p θ Equation 8
[0166] where the function THRESH(*, θ) binarizes the image through θ, φ(I) represents the feature map of image I and is normalized to the range of 0 - 255, and p φ and p θ represent the prior distribution. The influence of the Boolean map on visual attention is represented by the attention map A(B):
[0167]
[0168] where I is the input image, is the average attention map, which is used for subsequent processing to obtain a final differential superposition map S.
[0169] 2. Target sample update
[0170] Crop the alternative image blocks of the next frame around the target in the current frame, train the filter online in real time, then obtain the target position of the next frame, and update each other continuously for tracking. The strategy for judging the position information of the target is: the position corresponding to the maximum filter response value is the location of the target center.
[0171] However, this sample update method has the following two disadvantages:
[0172] First, regardless of whether the current image can be accurately tracked, its tracking results will be accumulated to the next frame image, affecting the subsequent tracking effect; there is no one-to-one correspondence between circular motion and the actual translation of the target. Circular motion is only an approximation of the actual translation in the image. In real tracking scenarios, especially in the presence of fast motion, occlusion, etc., it becomes unreliable. Using a single central Gaussian as the target response hinders the performance of the tracker and leads to irrecoverable drift.
[0173] Second, the existing technology tracking algorithms all adopt continuous and deterministic weight sample updates in sample updates, that is, the results obtained in each frame will affect the subsequent target, and the sample update weights will be directly determined in the initialization. Then, the influence of each frame image on the tracking result has also been determined. However, in actual tracking scenarios, the role of each frame image in sample training is different. Especially when the target is occluded, the target appearance features extracted at this time are quite different from the original target, basically the target background. Once misidentified as the target and added to the sample training, it will contaminate the training samples.
[0174] In the training stage of this application, the response value is considered in the weight of sample update, and the detection result in the tracking algorithm is utilized, and it does not simply adopt a binary strategy. Even when the target is occluded, the information therein is utilized to avoid target drift when the target is occluded. First, the peak sidelobe ratio is used to determine whether the tracking is effective, and then the influence of the current tracking result on the next frame tracking process is directly determined according to the maximum response value corresponding to the tracking result. Assuming that starting from the second frame image, the maximum response value of each image is stored in response_all, the update of the sample is as follows:
[0175]
[0176] where η is the learning factor, and respectively represent the Fourier transforms of the target matching templates of the current frame and the previous frame. The detection result is considered in the sample update. When the target is contaminated, the weight it occupies decreases, and the influence on the subsequent target also decreases accordingly.
[0177] 3. Tracking model update
[0178] Currently, the DCF-based tracking algorithms sacrifice tracking speed and real-time performance while pursuing tracking effects. Therefore, in this application, aiming at the target occlusion scenario during the tracking process, the tracking algorithm is improved in terms of model update to improve the speed.
[0179] When updating the model, a sparse update strategy is adopted. The model is updated every Ns frames, and the samples still need to be updated every frame. The reduced model update frequency saves time and can avoid the model drift problem, showing an improvement effect to a certain extent. However, Ns cannot be set too large, otherwise the model will not be able to keep up with the changes of the target. Finally, in this application, Ns is set to 6. Since more and more parameters are used to pursue higher accuracy, the appearance model and the motion model become more and more complex. For tracking, due to the lack of a large number of training samples, overfitting is likely to occur. Therefore, updating the model not every frame can effectively prevent model drift and save computing time at the same time.
[0180] III. Robust Scale Estimation of 2D Targets
[0181] The existing target tracking algorithms focus on estimating the position of the target. For the moving target in the consumer drone video, the scale changes while the target deforms. Only predicting the position of the target limits the tracking performance and it is difficult to achieve good tracking. Estimating the scale change of the target is beneficial to improving the tracking accuracy. After obtaining the position information of the target, this application trains a 2D filter to estimate the length and width of the target respectively, which can more robustly estimate the scale of the target.
[0182] (I) Scale Estimation Based on Perspective Projection Model
[0183] After knowing the motion speed of the drone motion platform, estimate the scale change of the target. The perspective imaging principle is as Figure 1 shown. In the figure, of is the focal length, on is the image distance, and om is the object distance. According to the optical imaging principle of the convex lens, when the object distance is much greater than the image distance (m > f), at this time, the focal length and the image distance are considered approximately equal, and the central imaging model approximately replaces the perspective imaging model. For details, see Figure 2 , where M is the point in the camera coordinate system, and m is the projection of point M in the image coordinate system. Let the vector expression of point M be M = (x M , y M , z M ) T , and the vector expression of point m be m = (x m , y m ) T . Under the perspective imaging model, the transformation relationship between the two points is as follows:
[0184]
[0185] The non - linear transformation formula from the camera coordinate system to the image coordinate system is as follows:
[0186]
[0187] If the distance between the moving platform and the target is known, the scale of the target in the phase plane can be estimated by combining the internal parameters of the camera.
[0188] (2) Two - dimensional scale estimation and tracking framework based on correlation locking
[0189] Design two consistent correlation filters, defined as the position target filter and the scale target filter, to achieve target tracking and scale transformation respectively. The former is used for the positioning of the target in the current frame, and the latter is used for the scale estimation of the target in the current frame. The two filters are relatively independent, and different feature types and feature calculation methods are selected for training and testing. A two - dimensional filter is adopted, assuming that the length and width of the target do not change in the same way. Even if the target undergoes large deformation, the target scale can still be accurately estimated.
[0190] As Figure 3 shown, the overall framework of the tracker with two - dimensional target robust scale estimation is divided into two parts: position prediction and scale size estimation. The specific steps are as follows: Training samples are obtained through dense sampling in the search area around the target and feature extraction is performed. The samples are subjected to Fourier transform, and the filter is trained using least - squares regression in the frequency domain; An online learning method is used to obtain a position kernel correlation filter. Then, the output value of the filter is mapped to the time domain through inverse discrete Fourier transform, and the coordinates of the maximum point in the response map are found. This point corresponds to the center position of the target in the sample; Then, the scale prediction process based on the correlation filter is similar to the process of predicting the position of the target. After obtaining the position of the target, according to the preset scale value, downsampling and upsampling are performed around the target to obtain a series of image patches with different scales; Then, these image patches are bilinearly interpolated to make the size of the image patches consistent with the designed scale model to obtain training samples; Next, feature extraction is performed. After obtaining the HOG features of the image patches, a least - squares classifier is trained to obtain a scale target tracker; In addition, after obtaining the features of the image patches, the features are windowed using a Hamming window to suppress the high - frequency noise caused by the image boundary in the frequency domain due to the use of a circulant matrix.
[0191] When applying the scale target tracker to a new image, calculate the score in the scale space, that is, the response value of the filter. The position of the maximum value in the response map of the scale target filter corresponds to the final scale of the target.
[0192] Construct a location target filter to obtain the location of the target to be tracked, construct a scale target filter, and use the target center location obtained by the location target filter to obtain the target scale, and transfer the final target information to the location target filter of the next frame of image. The two operate together, update the parameters online, and improve the performance of the entire tracker.
[0193] (III) Two-dimensional scale prediction evaluation strategy
[0194] Detect the scale change of the target by learning a two-dimensional scale correlation filter, and reasonably limit the target search area according to the scale change of the target to avoid unnecessary calculations.
[0195] The most important part of target scale estimation is to give specific alternative values of target scale transformation: in the scale prediction evaluation process, for the current image, the size of the target is P×R, and the size of the scale target filter is S×S. For each Extract image patches J of size in the two dimensions of the length and width of the target n×n , where a is the filter parameter factor. For targets with large scale changes, the alternative scales should be as different as possible. On the contrary, the alternative scales should change less.
[0196] (IV) Adaptive search window scale selection
[0197] Tracking based on a discriminant model uses binary classification to separate the target from the background and takes into account the motion information of the target. The position where the target appears in the next frame must be within the neighborhood centered on the target in the current frame. For the value of padding, it should be considered that: the relative displacement △p of the target center between two adjacent frames under the UAV motion platform is more variable than that in ordinary scenes, so that the target part in the next frame is sometimes not inside the search sub-window, resulting in the tracking result not corresponding to the target and tracking drift. The position of the target obtained by the algorithm in the next frame must be within the window of size windos_sz centered on the target in the current image. In this way, once the spatial context is determined and the actual target is not in this area, no matter how robust the previous classifier model is, it is impossible to detect the target, resulting in tracking failure. If the value of padding is too small, the target is not in the search box centered on the target in the previous frame of image; if the value of padding is too large, the retained background information increases, and it is easy to generate false detections when there are objects similar to the target in the background. In addition, too much data increases the computational amount.
[0198] To solve the situation where the target positions in two adjacent frames are far apart, the present application adopts an adaptive search window strategy to constrain the size of the search box according to the offset of the target positions in two adjacent frames;
[0199] For the tracker based on the correlation locking framework, when performing correlation operations, the scale of the filter is initially determined during initialization. The size of the search area is changed according to the offset of the target positions in two adjacent frames, and the size of the search window is adaptively selected by comparing the offsets of the target centers in two adjacent frames. A critical value is preset, and when the offset is greater than this critical value, the search window is enlarged.
[0200] (5) Scale prediction and estimation method flow
[0201] In the entire tracking algorithm after adding scale estimation, first, the position information is obtained using the position target filter, and then the current scale change of the target is obtained using the motion scale target filter. The kernel function involved in the algorithm is the Gaussian kernel function.
[0202] In the scale estimation module, HOG features are used. First, the features of the image are extracted to obtain the training samples f of the d-dimensional features of the image l , where l = {l, 2, …, d}, and each feature dimension corresponds to an association filter h l , minimizing the loss function:
[0203]
[0204] where g is the output of the filter designed for the training sample f, and λ ≥ 0 is the regularization parameter, controlling the structural error. Solving the equation gives the filter as:
[0205]
[0206] Adding a regularization factor, performing l 2 constraint on the filter, so that f has no non-zero terms, alleviating the problem of zeros appearing in the filter in the frequency domain and avoiding the phenomenon of the denominator being zero. Online solving the update problem of a d×d-dimensional linear equation system to obtain a robust approximation and updating the association filter H l t for the numerator A l t and the denominator B t :
[0207]
[0208] where η is the learning rate, and the final output of the filter, that is, the association score:
[0209]
[0210] The scale transformation corresponding to the maximum value in the correlation response map is the current scale size of the target.
[0211] When extracting features from the target alternative regions, the length and width of the target are respectively scaled to obtain alternative regions. Although the scale change of the target can be accurately estimated, the computational complexity is increased. The principal component feature transformation is used to reduce the dimension. The eigenvalues and eigenvectors of the covariance matrix of the high-dimensional data are calculated. The eigenvectors are the selected basis vectors, and the eigenvalues corresponding to the eigenvectors represent the projection of the original data on the vectors. The larger the eigenvalue, the more information of the original data is retained by the corresponding eigenvector, removing redundant information in the data and reducing the amount of data to be calculated.
[0212] In addition, for the scale target filter, the size of the target is concerned. To reduce the computational complexity and retain the overall information of the target, the target is downsampled to a certain size before calculating the target features, rather than processing the entire target.
[0213] IV. Experimental Results and Analysis
[0214] (I) Qualitative Result Analysis
[0215] The entire experiment was completed based on the Matlab platform simulation. It mainly solves the problem of scale estimation of the target when the target undergoes scale changes such as scale changes and out-of-plane rotations during the tracking process of a moving platform. The video sequences used include ships on the water, moving cars, and people on the water.
[0216] Based on the DSST as the basic algorithm, whether it is the position estimation or scale estimation of the target, the selected feature is HOG. Pay attention to the scale target filter of the algorithm. The parameter settings in the experiment are as follows: the padding value is 2, the number of scale alternative values is 17, the step size is 1.02, and the key frames in the image sequence are intercepted to visually judge the tracking effect of the algorithm.
[0217] For example Figure 4 , in the boat1 test video, the background is simple, and the target undergoes slow deformation and small scale changes. From the tracking results, it can be seen that the DSST tracking algorithm is affected by the initial shape of the target, and the tracking results cannot change with the deformation of the target. The algorithm of the present application is more accurate in scale estimation than DSST.
[0218] For example Figure 5 , in the boat2 test video, the target mainly undergoes huge deformations. The boat enters the field of view from the side and then circles, and then sails into the distance. From the tracking frames, the tracking results of the two are approximately the same, but obviously the algorithm proposed in the present application is closer to the ground truth and has better tracking performance.
[0219] For example Figure 6, in the boat3 test video, the background is simple and the target mainly undergoes deformation. It can be seen from the tracking results that the DSST tracking algorithm is affected by the initial shape of the target. After the perspective changes, the shape of the tracking result cannot change with the deformation of the target, thus affecting the entire tracking result. In this application, when the target becomes smaller, the tracking frame also becomes smaller, which is more accurate than the scale estimation of DSST.
[0220] Such as Figure 7 , in the car4 test video, the car drives from a roundabout into a straight road and is subject to interference such as occlusion, disappearance, and similar backgrounds. From the intercepted key frames, when the target is occluded until it completely disappears and then reappears in the field of view, the DSST algorithm fails to track, while the algorithm of this application can still continue to track the target. When there is a similar background around the target, the tracking is not affected either.
[0221] Such as Figure 8 , in the wakeboard5 test video, the target mainly undergoes large deformations and large scale changes. A person jumps into the sea from the shore and starts wakeboarding. The person moves quickly from side to side accompanied by actions such as jumping up, squatting down, and standing upright. At the same time, the person moves away from the shooting platform and gradually becomes smaller in the field of view. When the posture of the target itself changes, this application can be closer to the ground truth of the target, and the tracking result is also more accurate.
[0222] (2) Quantitative result analysis
[0223] Quantitative analysis is carried out on the experimental data. By using the OPE method, the evaluation parameters selected are distance accuracy and overlap accuracy. From the experimental results, the tracking effect of this application is superior to that of DSST in terms of both tracking accuracy and success rate. The algorithm using two-dimensional scale estimation can indeed more accurately estimate the scale of the target, thereby improving the overall tracking accuracy and enabling the algorithm to have better tracking performance.
Claims
1. A precise tracking method for moving objects in consumer drone videos, characterized in that, improvements are made in three aspects of deformation, occlusion, and scale based on the correlation lock tracker: First, a correlation lock tracking algorithm based on decision perception is proposed, including: one is a tracking framework based on decision perception; the second is feature correlation selection; the third is feature decision perception; the fourth is decision perception weight calculation; Second, a correlation lock tracking method based on difference superposition detection is proposed, including: one is rapid locking of moving objects; the second is correlation lock target tracking based on difference superposition detection; Third, a two-dimensional target robust scale estimation method is proposed, including: one is scale estimation based on the perspective projection model; the second is a two-dimensional scale estimation tracking framework based on correlation lock; the third is a two-dimensional scale prediction and evaluation strategy; the fourth is adaptive search window scale selection; the fifth is the process of scale prediction and estimation method; (1) For the drift caused by target deformation during tracking, different features are used for tracking respectively and fused at the decision level to obtain the final result. One method is to construct two correlation filters, using the HOG feature with target structure information and the CNs with target color information respectively. After calculating the tracking results, the maximum response value of the filter is used at the decision level to determine the weight of each filter for the final result. Another method is to use the correlation lock and a tracking framework based on the color histogram, and then fuse the results of the two to obtain the final tracking result; (2) For the target occlusion problem during tracking, a correlation lock tracking method based on difference superposition detection is proposed. First, the structural risk function in ridge regression is replaced with a regularization function, and components are added to better handle the occlusion problem. Then, the difference superposition feature is extracted within the search box and superimposed with the original sample to increase the difference between the target and the surrounding background, and at the same time reduce the potential risk of the circulant matrix. A method based on the detection result is adopted for sample update, and non-continuous update is used for model update to handle the target occlusion problem during the tracking process; (3) For the target scale change problem during the tracking process, a two-dimensional target robust scale estimation method is proposed. First, the rough scale estimation of the target is obtained using sensor data and the internal parameters of the shooting camera. A two-dimensional correlation filter is designed to estimate the target scale, and the length and width dimensions of the target are estimated separately. When the target deforms and scales change, the true size of the target in the field of view can be accurately estimated. In addition, an adaptive search window method is proposed to change the size of the search box according to the displacement of the target in two adjacent frames to prevent the target from moving too fast and moving outside the search window; The samples are downsampled on the scale target filter to reduce the amount of calculation data.
2. The precise tracking method for moving objects in consumer drone videos according to claim 1, characterized in that, Decision-aware Tracking Framework: Achieve the optimal tracking effect with the simplest possible model and a small number of parameters. Suppose in the t-th image I t the tracking problem is regarded as finding the most likely region p t from the set of candidate regions S t : Among them, T refers to performing a certain transformation on the image, f(T(I t , p); θ) represents the score of the rectangular box p in the image I t under the model parameters θ, and the model parameters satisfy minimizing the loss function: Among them, the loss function L(θ; χ t ) is related to the positions of the previous image and the target in the image, is the entire model parameter space, R(θ) and λ are regularization factors to prevent the model from being too complex and causing overfitting, and achieving effective tracking is transformed into the problem of selecting functions f and L; When the target is initialized, the position and scale information of the target are obtained. Finding the region most similar to the target in the alternative image sequence is to track the target. Judging whether two signals are correlated is to measure the similarity between them. When calculating the similarity between the alternative region sample and the target, an association operation is performed on the two. The loss function is selected to calculate the distance between two vectors. The tracking algorithm uses positive samples containing the target and negative samples that are all backgrounds, and uses ridge regression to construct a regressor to train the appearance model of the target.
3. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Feature association selection: Based on complementary features, multi-feature decision perception is performed on the moving target, and according to the different impacts of each feature on the tracking result, a framework is designed for fusion to obtain a more accurate and stable tracking effect; Among them, for HOG feature selection, the sample region is divided into several sub-regions, and then 32-dimensional features are extracted in each region. Except for the last dimension which is all 0 and not considered, the remaining 31-dimensional features are used. For texture features, the RFS filter bank of the fast anisotropic Gaussian filter is used, and this feature has a total of 8 dimensions.
4. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Feature decision perception: Two filters are trained with different input samples. One of the input images is the target box and a circle of background around it, which contains spatial context information, that is, the size is window_sz. To better distinguish the target and the background by superimposing differences, HOG features are selected. In addition, to eliminate the boundary effect, a cosine window function is added to the extracted features for processing; the other only contains the target, and the input image is smaller than the size of the target itself, that is, the size is sz. To make the tracker still have good robustness when the target undergoes large deformations, CNs features are used; The former adds background information to expand the search area of the target, and the latter only targets the target to improve accuracy; the two filters update the template online, and the tracking result is more accurate. By using multiple feature decision perceptions of the target, more target information is retained and the target is tracked by multiple means to better handle unexpected situations during the tracking process.
5. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Decision perception weight calculation: An additive model of multi-feature perception is adopted. Under the additive fusion strategy, the final results after n kinds of decision perceptions are additively fused. Additive fusion is insensitive to noise and increases the robustness of the tracking algorithm. According to the filter response value and the target position, the final target position is obtained as: Among them, res1 and res2 are the maximum response values of two filters, and p 1 and p 2 are the positions of the targets corresponding to the maximum response values respectively; The algorithm flow of using HOG and CNs decision perception: Input: Image sequence Initial position p of the target 0 ; Output: Position information of the target in each frame of the image Process 1: Initialization, t = 1; Process Two: Adopt I 0 and p 0 Initialize trackers Track1 and Track2; Process 3: for t = 1:T Process Four: Output p according to the previous frame t-1 Calculate the current search sub-window; Process 5: For the tracker Track1, HOG features are used, the target structure information is utilized, the search window contains spatial context information, and a Hanning window is used for window processing; Process 6: For the tracker Track2, CNs features are used, the global color information of the target is utilized, and the search window is the target itself without window processing; Process Seven: Obtain the response maps of the two trackers from the tracking model; Process Eight: Find the maximum response value in the response map, and this value corresponds to the most likely location of the target; Process Nine: Fusion is performed by an additive model with multi-feature perception to obtain the final tracking result p of the target 0 ; Output: Position information of the target in each frame of the image 6. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Fast Locking of Moving Target: The tracking algorithm based on fast locking of moving targets uses the filter W to associate with the alternative regions to estimate the target position. Among them, an image block that contains the target and has a size larger than the target is used as the search window, and then the search window is cyclically shifted to obtain different alternative windows. The actual displacement of the target is approximated through circular motion. Finally, a data matrix with a circular structure is constructed to obtain training samples; Suppose the kernel matrix K z is obtained by performing kernel correlation operations on the matching template and the search window samples, where the matching template X and the detection sample Z are circulant matrices obtained by shifting the vectors x and z, and K z is also a circulant matrix: K z = C(k xz ) Equation 4 where k xz is the kernel correlation operation of vector x and vector z. After introducing the kernel function in the least squares method, the regression function is f(z): f(z) = (K z )α Equation 5 α is the transformation coefficient, and performing DFT operations on both sides gives: Among them, z is the predicted image block of the search window including the target, and the position of the target is detected from it. x is the target model learned in the image. The right side of the equation is a linear combination of k multiplied by the transformation coefficient α. ⊙ is the element-by-element multiplication in the vector, and * is the conjugate. Performing the inverse DFT operation on Equation 6 gives the response matrix of the filter. The position corresponding to the maximum value in the matrix is the target position. The model update is as follows: where η is the learning factor, and are the Fourier transforms of the coefficient α for the current frame t and the previous frame t - 1 respectively, and represent the Fourier transforms of the target matching templates for the current frame and the previous frame respectively.
7. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Associative Locking Target Tracking Based on Difference Superposition Detection: Further improvements are made in two parts: sample training of the filter and update of filter parameters; 1. Difference Superposition Detection Adding a regularization parameter to prevent model overfitting and reducing the test error. Using sparse parameters to effectively handle occlusion. In the training sample stage, the difference superposition region is taken into account, and the parts that are easy to attract attention are preferentially processed, which is convenient for improving the tracking performance when the target is occluded. Focus on the background, remove the statistical redundancy of the input signal, and remove the redundant information of the target to obtain the salient target of the image. Use the log spectrum of the image. The log amplitude spectrum of an image minus the average log amplitude spectrum is the difference superposition part of the image; While considering the differential superposition information of the image and retaining the original information of the target, during the sample training process, Ω(u) = u + S BB (t), where u is the input image and S BB (t) is the salient region; After removing the redundant information of the image, deeper difference superposition detection is performed. Based on the instantaneous awareness of the scene, there is a Boolean mapping representation. According to the prior distribution and random threshold in the feature space, a series of Boolean maps are generated using the feature map of the input image: B i = THRESH(φ(I), θ) φ~p φ ,θ~p θ Equation 8 Among them, the function THRESH(*, θ) binarizes the image through θ, φ(I) represents the feature map of the image I, and is normalized between 0 and 255, p φ and p θ represents the prior distribution, and the influence of the boolean map on visual attention is represented by the attention map A(B): where I is the input image, is the average attention map, which is used for subsequent processing to obtain a final difference superimposed map S; 2. Target Sample Update In the training stage, the response value is considered in the weight of sample update, and the result detected by the tracking algorithm is utilized, rather than simply adopting a binary strategy. First, use the peak sidelobe ratio to determine whether the tracking is effective, and then directly determine the influence of the current tracking result on the next frame tracking process according to the maximum response value corresponding to the tracking result. Assume that starting from the second frame image, the maximum response value of each image is stored in response_all. The update for the sample is: where η is the learning factor, and respectively represent the Fourier transforms of the target matching templates of the current frame and the previous frame. The detection results are taken into account in the sample update. When the target is contaminated, the weight it occupies decreases, and the influence on subsequent targets also decreases accordingly; 3. Tracking Model Update When updating the model, a sparse update strategy is adopted. The model is updated every Ns frames, and the samples still need to be updated every frame. The model update frequency is reduced, saving time and avoiding the problem of model drift. Ns is taken as 6.
8. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Two-dimensional Scale Estimation Tracking Framework Based on Correlation Locking: Design two consistent correlation filters, defined as the position target filter and the scale target filter, to achieve target tracking and scale transformation respectively. The former is used for the localization of the target in the current frame, and the latter is used for the scale estimation of the target in the current frame. The two filters are relatively independent, and different feature types and feature calculation methods are selected for training and testing. A two-dimensional filter is adopted, and it is considered that the length and width of the target do not change in the same way. Even if the target undergoes large deformations, the target scale can still be accurately estimated; The overall framework of the tracker with two-dimensional target robust scale estimation is divided into two parts: position prediction and scale estimation. The specific steps are as follows: Training samples are obtained through dense sampling in the search area around the target and feature extraction is performed. The samples are subjected to Fourier transform, and the filter is trained using least squares regression in the frequency domain; An online learning method is used to obtain a position kernel correlation filter. Then, the output value of the filter is mapped to the time domain through inverse discrete Fourier transform, and the coordinates of the maximum point in the response map are found. This point corresponds to the center position of the target in the sample. Then, the scale prediction process based on the correlation filter is similar to the position process of the predicted target. After obtaining the position of the target, according to the preset scale value, downsampling and upsampling are performed around the target to obtain a series of image patches with different scales. Then, these image patches are bilinearly interpolated to make the size of the image patches consistent with the designed scale model to obtain training samples. Next, feature extraction is performed. After obtaining the HOG features of the image patches, a least squares classifier is trained to obtain a scale target tracker. In addition, after obtaining the features of the image patches, the features are windowed using a Hamming window to suppress the high-frequency noise caused by the image boundary in the frequency domain due to the use of a circulant matrix; When applying the scale target tracker to a new image, calculate the score in the scale space, that is, the response value of the filter. The position of the maximum value in the response map of the scale target filter corresponds to the final scale of the target; Construct a position target filter to obtain the position of the target to be tracked, construct a scale target filter and use the target center position obtained by the position target filter to obtain the target scale, and transfer the final target information to the position target filter of the next frame image. The two work together and update the parameters online to improve the performance of the entire tracker.
9. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Two-dimensional Scale Prediction Evaluation Strategy: Detect the scale change of the target by learning a two-dimensional scale correlation filter, and reasonably limit the target search area according to the scale change of the target; The most important part of target scale estimation is to give specific alternative values for target scale transformation: during the scale prediction and evaluation process, for the current image, the size of the target is P×R, and the size of the scale target filter is S×S. For each image patches J of size are extracted separately in the two dimensions of the length and width of the target n×n , where a is the filter parameter factor. For targets with large scale changes, the alternative scales should be as different as possible; conversely, the alternative scales should change less. Adaptive Search Window Scale Selection: The size of the search box is constrained according to the offset of the target positions in two adjacent frames. Based on the tracker of the association locking framework, when performing the association operation, the scale of the filter is initially determined during initialization. The size of the search area is changed according to the offset of the target positions in two adjacent frames. The size of the search window is adaptively selected by comparing the offset of the target centers in two adjacent frames. A critical value is preset. When the offset is greater than this critical value, the search window is enlarged.
10. The method for accurately tracking a moving target in a consumer drone video according to claim 1, characterized in that Scale Prediction and Estimation Method Process: In the entire tracking algorithm after adding scale estimation, first, the position information is obtained using the position target filter, and then the current scale change of the target is obtained using the motion scale target filter. The kernel function involved in the algorithm uses the Gaussian kernel function; In the scale estimation module, HOG features are adopted. First, feature extraction is performed on the image to obtain the training sample f of the d-dimensional features of the image l , where l = {l, 2, …, d}, and each feature dimension corresponds to an associated filter h l , minimizing the loss function: where g is the output of the filter designed for the corresponding training sample f, and λ≥0 is the regularization parameter that controls the structural error. Solving the equation gives the filter as: Add a regularization factor and perform an l 2 constraint on the filter so that f has no non-zero terms, alleviating the problem of zeros in the frequency domain of the filter and avoiding the phenomenon of a zero denominator. Solve the update problem of a d×d-dimensional linear equation system online to obtain a robust approximation and update the associated filter H l t for the numerator A l t and the denominator B t : where η is the learning rate, and the output of the final filter, that is, the association score: The scale transformation corresponding to the maximum value in the correlation response map is the current scale size of the target; When extracting features from the target candidate regions, the length and width of the target are respectively scaled to obtain the candidate regions. Although the scale change of the target can be accurately estimated, the computational amount is increased. The principal component feature transformation is used for dimensionality reduction. The eigenvalues and eigenvectors of the covariance matrix of the high-dimensional data are obtained. The eigenvectors are the selected basis vectors. The eigenvalues corresponding to the eigenvectors represent the projection of the original data on the vectors. The larger the eigenvalue, the more information of the original data is retained by the corresponding eigenvector. The redundant information of the data is removed to reduce the amount of data for calculation; In addition, on the scale target filter, the focus is on the size of the target. To reduce the computational amount and retain the overall information of the target, the target is downsampled to a certain size before calculating the target features, rather than processing the entire target.
Citation Information
Patent Citations
System and method for identifying target objects
CN107851308A
Unmanned ship intelligent decision-making method and system
CN110782481A