Target loss recapture method based on correlation filtering

By adding adaptive scale change judgment and SURF feature matching based on the relevant filtering algorithm, combined with weighted judgment logic and recognition algorithm, the problem of low tracking success rate and accuracy during fast moving targets, long-term occlusion and target size changes is solved, and efficient target loss recapture and long-term stable tracking are achieved.

CN120125614APending Publication Date: 2025-06-10CHONGQING JIANSHE IND GRP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510174725.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has low success rate and accuracy in tracking when facing complex situations such as fast moving targets, long-term occlusions, and continuous changes in target size.

Method used

The target loss recapture method based on correlation filtering is used to calculate the average value of the highest peak value of the correlation filter response and its average peak correlation energy, and combine the weighted judgment logic to judge whether the target is lost; use the cosine similarity value to determine whether the target reappears; add scale change judgment and SURF feature matching, reposition and capture the target position; when the target is about to be lost, template updates are paused, keyframe features are retained, and recaptured using recognition algorithms or SURF feature matching.

Benefits of technology

It improves the ability to adapt to target scale changes, enhances the success rate and accuracy of target loss recapture, and can effectively track the target under severe shaking and long-term occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125614A_ABST
    Figure CN120125614A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target tracking, in particular to a target loss and recapture method based on correlation filtering, and the method comprises the steps: judging the target loss and recapture conditions through a correlation filtering algorithm and weighted logic; when the target appears again, judging whether the scale changes or not, and then increasing or reducing the corresponding step length to change the scale; the method comprises the following steps of: carrying out target repositioning by adopting an SURF (Speeded Up Robust Features), carrying out target detection by using the extracted SURF after a locked target is lost, and when the target appears again, re-determining the position of the target through feature matching; when the target is about to be lost, template updating is suspended, key frame features are reserved, and a recognition algorithm or SURF feature matching is utilized to recapture the target position; and the next frame of video stream is read for continuous tracking, and if the video stream reading is finished or the tracking fails, the tracking is stopped, so that the technical problems of low tracking success rate and low tracking accuracy in the prior art under the complex conditions of a fast moving target, long-time shielding, continuous change of the target size and the like are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and in particular to a method for re-capturing a lost target based on correlation filtering. Background Art

[0002] Target tracking technology is an algorithm that uses the information of a known template to determine the target position and motion state in a captured video sequence based on the temporal and spatial relationships of the target. Target tracking is an important issue in the field of computer vision. There are various classification methods for visual target tracking. Currently, the more popular and widely recognized classification method is to divide the tracking methods into generative methods and discriminative methods according to the expression strategy of the target appearance model. The generative method is to model the target area in the current frame, and find the area most similar to the model in the next frame to predict the position of the target. Currently, target tracking methods mainly focus on correlation-filtering-based target tracking methods and deep-learning-based target tracking methods. The discriminative method regards the tracking problem as a classification problem between the target and the background. By learning the target and the background, the predicted area is used as the position where the target is located. The correlation-filtering-based method is an important development direction of the discriminative method, and its main feature is fast operation speed, meeting the real-time requirement.

[0003] The first to apply the correlation filter to the tracking field were Bolme et al., who proposed a Minimum Output Sum of Squared Errors (MOSSE) correlation filter tracking method. This algorithm only uses a sample image of the target area to train the target appearance model. By using the discrete Fourier transform, the similarity calculation between the target model and the candidate area is transformed into the frequency domain, significantly improving the running speed of the tracking algorithm. The tracking speed reached 615 frame / s, but the tracking accuracy and success rate still need to be further improved. Henriques et al. proposed the Circulant Structure of Tracking with Kernels (CSK) algorithm. This algorithm constructs a large number of training samples by circularly shifting the reference training samples to train the classifier. At the same time, a large number of samples to be tested are constructed in the same way for the target detection process. By using the property that the circulant matrix can be Fourier diagonalized, the calculations of classifier training and target detection are transformed into the frequency domain to achieve fast calculation. Henriques et al. proposed the Kernel Correlation Filter (KCF) tracking algorithm based on CSK. It uses the Histogram of Oriented Gradients (HOG) feature to replace the original grayscale value feature, expands the correlation filter from a single channel to multiple channels, improving the tracking accuracy and having good real-time performance.

[0004] In the face of complex situations such as fast-moving targets, long-term occlusion, and continuous changes in target size, the success rate and accuracy of tracking using the above methods are relatively low. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for re-capturing a lost target based on correlation filtering, aiming to solve the technical problem that in the face of complex situations such as fast-moving targets, long-term occlusion, and continuous changes in target size in the prior art, the success rate and accuracy of tracking are relatively low.

[0006] To achieve the above purpose, a method for re-capturing a lost target based on correlation filtering adopted by the present invention includes the following steps:

[0007] Based on the correlation filter algorithm, by calculating the average value of the highest peak of the correlation filter response and its average peak correlation energy for each frame, and combining the weighted judgment logic to determine whether the target is lost;

[0008] Use the cosine similarity value (calculate the cosine value of the vector angle) to determine whether the target has reappeared;

[0009] When the target reappears, add the comparison between two large-scale and two small-scale peaks and the peaks of the same scale to determine whether the scale changes, and then increase or decrease the corresponding step size for scale change;

[0010] Use SURF features to re-locate the lost target. After locking the target loss, use the extracted SURF features for target detection. When the target appears again, re-determine the target position through the matching of SURF features;

[0011] When the target is about to be lost, pause the template update, retain the key-frame features, and use the recognition algorithm or SURF feature matching, combined with color and HOG features, to re-capture the target position;

[0012] Read the next frame of the video stream to continue tracking. If the video stream reading ends or the tracking fails, stop tracking.

[0013] Among them, based on the correlation filtering algorithm, the specific method of judging whether the target is lost by calculating the average value of the highest peak of the correlation filtering response and the average peak correlation energy of each frame, combined with the weighted judgment logic, is as follows:

[0014] The average value of the highest peak S of the correlation filtering response calculated when the tracker updates the target position for each frame max and the average value of the average peak correlation energy APCE, and judge whether the target is lost by weighting the highest peak of the correlation filtering response and the average peak correlation energy of the current frame with the average values of S max and APCE of all previous frames;

[0015] The calculation formula of APCE is:

[0016]

[0017] In the formula: w and h are the row value and column value of the response matrix respectively; F i,j is the response value at the (i, j) position of the response matrix; mean(·) is to calculate the average; F max is the highest response value; F min is the lowest response value;

[0018] The occlusion criterion can reflect the oscillation degree of the response map. When APCE suddenly decreases, it means that the target is severely occluded or the target is about to be lost; here, the weighted values of the two maximum peaks F max of APCE are used to increase the accuracy of judgment. The occlusion criterion is as follows:

[0019] curr apce <A1×mean apce ;

[0020] curr_F max<B1×mean_F max ;

[0021] Where: A1 and B1 are adjustment factors, A1 = 0.62, B1 = 0.65; curr_apce is the APCE value of the current frame; mean_apce is the average value of APCE of all previous frames; curr_F max is the maximum response value of the current frame; mean_F max is the average value of the maximum response values of all previous frames. If both formulas hold, the judgment indicates that the target is about to be lost or partially occluded;

[0022] curr_apce < A2×mean_apce;

[0023] curr_F max <B2×mean_F max ;

[0024] A2 and B2 are adjustment factors, A1 = 0.45, B1 = 0.45; When both formulas hold, the judgment indicates that the target is completely lost or severely occluded.

[0025] Among them, the specific method of using the cosine similarity value to judge whether the target has reappeared is as follows:

[0026] Cosine similarity measures the similarity between two vectors by calculating the cosine value of the included angle between them. For two n-dimensional vectors S1 and S2, S1 = [a1, a2,..., an],

[0027] B = [b1, b2,..., bn], then the similarity calculation formula between A and B is as follows:

[0028]

[0029] Among them, the closer the cosine value is to 1, the closer the included angle is to 0°, the more similar the two vectors are, and the more similar the two images represented by the vectors are; Set a threshold U. When the calculated cosine value is greater than U, it means that the two frames of images are similar enough and the occluded target has reappeared.

[0030] Among them, by adding the comparison of two large-scale and two small-scale peaks with the peaks of the same scale to judge whether the scale changes, and then increasing or decreasing the corresponding step size for scale change. The specific method is as follows:

[0031] Judge whether the scale changes by adding the comparison of two larger-scale and two smaller-scale peaks. If the scale changes, the window size is corrected by the corresponding step size; if the scale does not change, the window size is not corrected. The scale peak comparison formula is:

[0032] F max(The first large scale)*scale_weight1>F max (Same scale);

[0033] F max (The second large scale)*scale_weight2>F max (Same scale);

[0034] F max (The first small scale)*scale_weight1>F max (Same scale);

[0035] F max (The second small scale)*scale_weight2>F max (Same scale);

[0036] Among them, scale_weight1 and scale_weight2 (0.9 ≤ scale_weight2 < scale_weight1 ≤ 0.99) are attenuation coefficients. If the attenuation ratio is larger than the same scale, it is considered as the target. The scale corresponding to the maximum value among the five is the current optimal scale. When the target expands or shrinks, the detection scale changes with the size of the target (in this way, both the target feature information can be retained to the greatest extent and the interference of background information on the target template can be suppressed. Through this method, the scale adaptation ability of the algorithm can be improved).

[0037] Among them, when the target is about to be lost, the template update is paused, and the key frame features are retained. The specific method for recapturing the target position by using the recognition algorithm or SURF feature matching, combined with color and HOG features is as follows:

[0038] For the target that can be recognized, through the YOLOV5 detection algorithm, detect all targets of the same type as the tracked target. Calculate the template image of the target found and retained one frame before the loss as the reference image. Calculate the similarity calculated by the combined color information CN of the recognized target and the similarity calculated by the features obtained by the HOG algorithm to perform template matching to calculate the highest peak value F of the fusion response map of the target max , calculate the scores of the appearance model in each recognition area. After the highest peak value of the fusion response map meets the set threshold C, it is the new position of the target;

[0039] For the target that cannot be recognized, use the feature matching algorithm for recapture.

[0040] Among them, when recapturing the target that can be recognized, the set threshold C is a hyperparameter, 0.4 ≤ C ≤ 0.6.

[0041] A method for re-capturing a lost target based on correlation filtering according to the present invention. First, based on the correlation filtering algorithm, by calculating the average value of the highest peak value of the correlation filtering response and the average peak correlation energy of each frame, and combining a weighted judgment logic to judge whether the target is lost; then, using the cosine similarity value to judge whether the target has reappeared; when the target reappears, add the comparison of two large-scale and two small-scale peaks with the same-scale peaks to judge whether the scale changes, and then increase or decrease the corresponding step size for scale change; adopt SURF features to re-locate the lost target. After locking the lost target, use the extracted SURF features to detect the target. When the target appears again, re-determine the target position through the matching of SURF features; when the target is about to be lost, pause the template update, retain the key frame features, and use the recognition algorithm or SURF feature matching, combined with color and HOG features, to re-capture the target position; finally, read the next frame of the video stream to continue tracking. If the video stream reading ends or the tracking fails, stop tracking. In this way, the technical problems in the prior art that the success rate and accuracy of tracking are relatively low in the face of complex situations such as fast-moving targets, long-term occlusion, and continuous change of target size are solved.

[0042] The present invention improves the scale change of target tracking by adding an adaptive scale change judgment on the basis of the correlation filtering algorithm, which can enhance the adaptive change ability of the algorithm to the target scale; by adding the target loss judgment of the algorithm and judging whether the target is occluded or disappeared and then reappears, after judging that the target reappears, introduce an identification module or predict the position where the target appears based on SURF feature matching, and then perform template matching by calculating the similarity calculated by combining the color information CN of the identified target and the similarity calculated by the features obtained by the HOG algorithm, and perform HOG and color histogram fusion feature template matching to determine the position where the target appears, re-capture the target position and re-initialize the target tracking, and continue tracking, so as to achieve long-term stable tracking of the target, and the re-capture time is fast, the success rate is high, and the tracking accuracy is high. It can effectively improve the success rate and accuracy of re-capture under the violent shaking of the target and long-term target occlusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0044] Figure 1 It is a re-capture schematic diagram of the method for re-capturing a lost target based on correlation filtering according to the present invention. Detailed implementation manners

[0045] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0046] Please refer to Figure 1 , Figure 1 which is a redetection schematic diagram of the target loss redetection method based on correlation filtering of the present invention.

[0047] The present invention provides a target loss redetection method based on correlation filtering, including the following steps:

[0048] S1. Manually select a tracking target, determine a sampling area of the tracking target, extract a feature map of the sampling area of the tracking target, and initialize tracker parameters;

[0049] S2. Based on the correlation filtering algorithm, by calculating the average value of the highest peak of the correlation filtering response and the average peak correlation energy of each frame, and combining a weighted judgment logic to judge whether the target is lost;

[0050] For this specific implementation manner, based on the correlation filtering algorithm, the specific manner of judging whether the target is lost by calculating the average value of the highest peak of the correlation filtering response and the average peak correlation energy of each frame and combining the weighted judgment logic is as follows:

[0051] The average value of the highest peak S max of the correlation filtering response calculated when the tracker updates the target position for each frame and the average value of the average peak correlation energy APCE are used to judge whether the target is lost by weighting the highest peak of the correlation filtering response and the average peak correlation energy of the current frame with the average values of S max and APCE of all previous frames;

[0052] The calculation formula of APCE is:

[0053]

[0054] where: w and h are the row value and column value of the response matrix respectively; F i,j is the response value at the (i, j) position of the response matrix; mean(·) is to calculate the average; F max is the highest response value; F min is the lowest response value;

[0055] The occlusion criterion can reflect the oscillation degree of the response map. When APCE suddenly decreases, it means that the target is severely occluded or the target is about to be lost; here, the weighted values of the two maximum peaks F max of APCE are used to increase the accuracy of the judgment. The occlusion criterion is as follows:

[0056] Curr apce <A1×mean apce ;

[0057] curr_F max <B1×mean_F max ;

[0058] Where: A1 and B1 are adjustment factors, A1 = 0.62, B1 = 0.65; curr_apce is the APCE value of the current frame; mean_apce is the average value of the APCE of all previous frames; curr_F max is the maximum response value of the current frame; mean_F max is the average value of the maximum response values of all previous frames. If both formulas hold, the judgment indicates that the target is about to be lost or partially occluded;

[0059] curr_apce < A2×mean_apce;

[0060] curr_F max <B2×mean_F max ;

[0061] A2 and B2 are adjustment factors, A1 = 0.45, B1 = 0.45; When both formulas hold, the judgment indicates that the target is completely lost or severely occluded.

[0062] S3. Use the cosine similarity value (calculate the cosine value of the vector angle) to judge whether the target has reappeared;

[0063] For this specific embodiment, the cosine similarity is used to measure the similarity between two vectors by calculating the cosine value of the angle between them. S1 and S2 are two n-dimensional vectors, S1 = [a1, a2,..., a n],

[0064] B = [b1, b2,..., b n], then the similarity calculation formula of A and B is as follows:

[0065]

[0066] Among them, the closer the cosine value is to 1, the closer the angle is to 0°, the more similar the two vectors are, and the more similar the two images represented by the vectors are; Set a threshold U. When the calculated cosine value is greater than U, it means that the two frames of images are similar enough and the occluded target has reappeared.

[0067] S4. When the target reappears, add the comparison of two large-scale and two small-scale peaks with the same-scale peaks to judge whether the scale changes, and then increase or decrease the corresponding step size for scale change;

[0068] For this specific embodiment, by adding two larger-scale and two smaller-scale peak comparisons to determine whether scale changes occur, if scale changes occur, the window size is corrected by the corresponding step length, and if the scale does not change, the window size is not corrected. The scale peak comparison formula is as follows:

[0069] F max (The first large scale)*scale_weight1 > F max (Same scale);

[0070] F max (The second large scale)*scale_weight2 > F max (Same scale);

[0071] F max (The first small scale)*scale_weight1 > F max (Same scale);

[0072] F max (The second small scale)*scale_weight2 > F max (Same scale);

[0073] Among them, scale_weight1 and scale_weight2 (0.9 ≤ scale_weight2 < scale_weight1 ≤ 0.99) are attenuation coefficients. If the attenuation ratio is larger than the same scale, it is considered the target. The scale corresponding to the maximum value among the five is the current best scale. When the target expands or shrinks, the detection scale changes with the size of the target (in this way, both the target feature information can be retained to the greatest extent and the interference of background information on the target template can be suppressed, and the scale adaptation ability of the algorithm can be improved through this method).

[0074] S5. Use SURF features to re-locate the lost target. After locking the lost target, use the extracted SURF features to detect the target. When the target appears again, re-determine the target position through the matching of SURF features;

[0075] For this specific embodiment,

[0076] S6. When the target is about to be lost, pause template update, retain the key frame features, and use the recognition algorithm or SURF feature matching, combined with color and HOG features, to re-capture the target position;

[0077] For this specific embodiment, for the recognized targets, through the YOLOV5 detection algorithm, all targets of the same type as the tracked target are detected. For the found targets, the template image retained one frame before the loss is calculated as the reference image. The fusion response map highest peak value F of the target is calculated through template matching by calculating the similarity calculated by combining the color information CN of the recognized target and the similarity calculated by the features obtained through the HOG algorithm. max , calculate the scores of the appearance model in each recognition area. After the highest peak value of the fusion response map meets the set threshold C, it is the new position of the target; the set threshold C is a hyperparameter, 0.4 ≤ C ≤ 0.6.

[0078] For the targets that cannot be recognized, the feature matching algorithm is used for recapture. The specific steps are as follows:

[0079] For the targets that cannot be recognized, after the target is lost, the extracted SURF features are used for target detection; the SURF features have a faster running speed compared to the SIFT features. Assume that the set of feature points successfully matched within the feature range is S = {(x 1 , y 1 ), (x 2 , y 2 ),....., (x n , y n )}, (x, y) are the coordinates of the corresponding points, and the position coordinates after SURF feature matching are When it is determined that the target reappears, the new position of the target is repositioned through the matching of SURF features, and the score (the highest peak value F of the fusion response map max ) in the new position obtained after SURF feature matching of the appearance model is calculated. After meeting the set threshold score C, it is determined as the new position of the target and the tracker is re-initialized for tracking. If the threshold score is not met within three frames, the recapture range is expanded.

[0080] The new recapture range is centered on the new target position after feature matching to the new target position. Assume that the size of the selected target box is a×b, and the set sampling window is four times the target box. Then the determined new target position box is (a×b), and the recapture search range is 2a×2b (that is, the recapture search range is four times the target area, and the recapture search range can be determined according to the target movement speed), where a and b represent the number of pixel points. As shown in the recapture schematic diagram: centered on the determined new target position (attached Figure 1Take the point at the center point of label 5 as the starting point for the search. When searching, consider the situation where the target may be between the two search boxes, so there is an overlap in the search range. Therefore, divide the 2a×2b re-capture search area into 9 search areas, each with a size of a×b. When searching, use these 9 points in the figure as the centers respectively to ensure that the target can be located even if it is between the two windows.

[0081] Let the Figure 1 The pixel coordinates of the center point of label 5 are (m, n), and the coordinates of the other 8 center points are (m - b / 2, n - a / 2), (m - b / 2, n), (m - b / 2, n + a / 2), (m, n - a / 2), (m, n + a / 2), (m + b / 2, n - a / 2), (m + b / 2, n), (m + b / 2, n + a / 2). Calculate the highest peak value F of the fusion response map in each of the nine windows max 2), (m, n - a / 2), (m, n + a / 2), (m + b / 2, n - a / 2), (m + b / 2, n), (m + b / 2, n + a / 2). Calculate the highest peak value F of the fusion response map in each of the nine windows max , and find the maximum peak value that is greater than the set threshold C, which is the target position sought.

[0082] After the predicted target appears, if the maximum matching peak value found is less than C, then continue to the next frame and return to template matching, determine the position and then perform the matching again, and continue for m frames (m is a hyperparameter. After data testing, according to the target speed, 10 < m < 20 can be taken, that is, if the target is not recaptured after 10 to 20 frames of re-capture), if no target meeting the conditions is found, it is determined that the tracking fails

[0083] S7. Read the next frame of the video stream and continue tracking. If the reading of the video stream ends or the tracking fails, stop tracking.

[0084] The HOG algorithm is as follows:

[0085] Extract the HOG feature vector of the target in the current frame, calculate the gradients in the horizontal and vertical directions of the image, use [-1, 0, 1] to calculate the horizontal gradient, and use its transpose to calculate the vertical gradient. The gradient of the pixel point (x, y) in the image is:

[0086] G x (x, y) = H(x + 1, y) - H(x - 1, y);

[0087] G t (x, y) = H(x, y + 1) - H(x, y - 1);

[0088] In the formula, G x (x, y) represents the horizontal gradient of the pixel point (x, y), and G y (x, y) represents the vertical gradient of the pixel point (x, y). Through G x (x, y) and G y(x, y) calculates the gradient magnitude and direction of this pixel point:

[0089]

[0090] In the formula, G(x, y) is the gradient magnitude and θ(x, y) is the gradient direction. The local image gradient information is statistically analyzed and quantified to obtain the feature description vector of the local image. Calculating the feature vector facilitates the matching with the target features of the previous frame.

[0091] Using a method for re-capturing a lost target based on correlation filtering according to the present invention, first, based on the correlation filtering algorithm, by calculating the average value of the highest peak value of the correlation filtering response and its average peak correlation energy for each frame, and combining the weighted judgment logic to judge whether the target is lost; then, using the cosine similarity value to judge whether the target has reappeared; when the target reappears, adding the comparison of two large-scale and two small-scale peaks with the peaks of the same scale to judge whether the scale has changed, and then increasing or decreasing the corresponding step size for scale change; using SURF features to re-locate the lost target. After locking the lost target, using the extracted SURF features to detect the target. When the target appears again, re-determining the target position through the matching of SURF features; when the target is about to be lost, pausing the template update, retaining the features of the key frames, and using the recognition algorithm or SURF feature matching, combining color and HOG features, to re-capture the target position; finally, reading the next frame of the video stream to continue tracking. If the reading of the video stream ends or the tracking fails, the tracking is stopped. In this way, the technical problems in the prior art that the success rate and accuracy of tracking are relatively low in the face of complex situations such as fast-moving targets, long-term occlusion, and continuous change of target size are solved.

[0092] The present invention improves the scale change of target tracking by adding an adaptive scale change judgment on the basis of the correlation filtering algorithm, which can enhance the adaptive change ability of the algorithm to the target scale; by adding the judgment of target loss in the algorithm and judging whether the target reappears after being occluded or disappeared, after judging that the target reappears, introducing an identification module or predicting the position where the target appears based on SURF feature matching for the target, and then performing template matching by calculating the similarity calculated by combining the color information CN of the recognized target and the similarity calculated by the features obtained by the HOG algorithm, and performing HOG and color histogram fusion feature template matching to determine the position where the target appears, re-capturing the target position to re-initialize the tracking of the target, and continuing to track, so as to realize the long-term stable tracking of the target, with fast re-capture time, high success rate, and high tracking accuracy, and can effectively improve the success rate and accuracy of re-capture under the severe shaking of the target and long-term target occlusion.

[0093] The above-disclosed is only a preferred embodiment of the present invention. Of course, the scope of rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A target loss recapture method based on correlation filtering, characterized in that: The steps include: Based on the correlation filtering algorithm, the average value of the highest peak value of the correlation filtering response and its average peak correlation energy of each frame is calculated, and combined with the weighted judgment logic to determine whether the target is lost; Use the cosine similarity value to determine whether the target has reappeared; When the target reappears, add two large-scale and two small-scale peaks and compare them with the peaks of the same scale to determine whether the scale has changed, and then increase or decrease the corresponding step size to change the scale; SURF features are used to relocate the lost target. After the target is lost, the extracted SURF features are used to detect the target. When the target appears again, the target position is re-determined by matching the SURF features. When the target is about to be lost, the template update is suspended, the key frame features are retained, and the recognition algorithm or SURF feature matching is used to combine color and HOG features to recapture the target position; Read the next frame of video stream to continue tracking. If the video stream reading ends or tracking fails, stop tracking.

2. The target loss recapture method based on correlation filtering as claimed in claim 1, characterized in that: Based on the correlation filtering algorithm, by calculating the average value of the highest peak value of the correlation filtering response and its average peak correlation energy of each frame, combined with the weighted judgment logic, the specific method of judging whether the target is lost is as follows: When the tracker updates the target position in each frame, the highest peak value S of the correlation filter response is calculated. max The average value of the average peak correlation energy APCE is calculated by comparing the highest peak value and the average peak correlation energy of the correlation filter response of the current frame with the S of all previous frames. max The average value of APCE is weighted to determine whether the target is lost; The APCE calculation formula is: Where: w, h are the row and column values ​​of the response matrix respectively; F i,j is the response value at (i, j) in the response matrix; mean(·) is the average; F max is the highest response value; F min is the minimum response value; The occlusion criterion can reflect the oscillation degree of the response graph. When APCE suddenly decreases, it means that the target is severely blocked or the target is about to be lost. Here, the maximum APCE peak F is used. max The weighting of the two values ​​increases the accuracy of the judgment. The occlusion criterion is as follows: curr apce <A1×mean apce ; curr_F max <B1×mean_F max ; Where: A1 and B1 are adjustment factors, A1 = 0.62, B1 = 0.65; curr_apce is the APCE value of the current frame; mean_apce is the average APCE value of all previous frames; curr_F max is the maximum response value of the current frame; mean_F max is the average value of the maximum response value of all previous frames. If both formulas are true at the same time, a judgment is made, indicating that the target is about to be lost or partially blocked; curr_apce<A2×mean_apce; curr_F max <B2×mean_F max ; A2 and B2 are adjustment factors, A1 = 0.45, B1 = 0.45; when the two formulas are established, a judgment is made that the target is completely lost or severely blocked.

3. The target loss recapture method based on correlation filtering as claimed in claim 2, characterized in that: The specific method of using cosine similarity to determine whether the target has reappeared is as follows: Cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. S1 and S2 are two n-dimensional vectors, S1 = [a1, a2, ..., an], B = [b1, b2, ..., bn], then the similarity calculation formula between A and B is as follows: The closer the cosine value is to 1, the closer the angle is to 0°, the more similar the two vectors are, and the more similar the two images represented by the vectors are. A threshold U is set. When the calculated cosine value is greater than U, it means that the two frames of images are similar enough and the occluded target reappears.

4. The target loss recapture method based on correlation filtering as claimed in claim 3, characterized in that: Add two large-scale and two small-scale peaks to compare with the peaks of the same scale to determine whether the scale has changed, and then increase or decrease the corresponding step size to change the scale. The specific method is: By adding two larger scales and two smaller scales to compare the peak values, we can determine whether the scale changes. If the scale changes by the corresponding step size, the window size is corrected. If the scale does not change, the window size is not corrected. The scale peak comparison formula is: F max (first large scale)*scale_weight1>F max (same scale); F max (Second large scale)*scale_weight2>F max (same scale); F max (first small scale)*scale_weight1>F max (same scale); F max (Second small scale)*scale_weight2>F max (same scale); Among them, scale_weight1, scale_weight2 (0.9=<scale_weight2<scale_weight1≤0.99) are reduction coefficients. A target is considered to be a target if the reduction ratio is larger than that of the same scale. The scale corresponding to the maximum value of the five is the current optimal scale. When the target is enlarged or reduced, the detection scale changes with the size of the target.

5. The target loss recapture method based on correlation filtering as claimed in claim 4, characterized in that: When the target is about to be lost, the template update is suspended, the key frame features are retained, and the recognition algorithm or SURF feature matching is used to combine the color and HOG features to recapture the target position. The specific method is as follows: For the targets that can be identified, the YOLOV5 detection algorithm is used to detect all targets of the same type as the tracked target. The template image of the found target is calculated and retained as a reference image before the loss. The template matching is performed by calculating the similarity of the combined color information CN of the identified target and the similarity of the feature calculation obtained by the HOG algorithm to calculate the highest peak F of the fusion response image of the target. max , calculate the score of the appearance model in each recognition area, and after the highest value of the peak of the fusion response map meets the set threshold C, it is the new position of the target; For targets that cannot be identified, feature matching algorithms are used for recapture.

6. The target loss recapture method based on correlation filtering as claimed in claim 5, characterized in that: When recapture is performed on an identifiable target, the threshold C is set as a hyperparameter, 0.4≤C≤0.6.

Citation Information

Cited By

  • Dynamic target tracking device and method based on visual information of cluster unmanned aerial vehicle

    CN120782816A