Complex scene single target tracking method, device, system and storage medium based on Staple algorithm
By introducing background weight histograms, similar target re-identification, loss judgment, feature adaptive fusion and template update strategies in the Staple algorithm, the problem of insufficient target tracking accuracy and robustness in complex scenarios is solved, and more efficient target tracking performance is achieved.
Patent Information
- Application Number
- CN202410826475.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-06-25
AI Technical Summary
When existing target tracking algorithms deal with complex scenarios, such as target deformation, occlusion, similar target interference and complex backgrounds, there are problems of insufficient accuracy and robustness.
A complex single-object tracking method based on Staple algorithm improves tracking performance through background weight histogram, similar target re-identification, loss determination, feature adaptive fusion and template update strategies.
It significantly improves the tracking accuracy and robustness of the target in complex scenarios, reduces error accumulation, and increases the probability that the target will be successfully detected and reconstructed after loss or occlusion.
Smart Images

Figure CN118864524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target tracking technology, and in particular to a complex scene single target tracking method, device, system and storage medium based on a Staple algorithm. Background Art
[0002] Object tracking is an ongoing research hotspot in computer vision and is crucial to many key applications such as security monitoring, military reconnaissance, and autonomous driving. With the expansion of application scenarios, object tracking algorithms face higher performance requirements, especially when dealing with complex scenarios such as occlusion, object deformation, interference from similar objects, and complex backgrounds. Existing algorithms need to show higher accuracy and robustness.
[0003] At present, single target tracking algorithms are mainly divided into two categories: traditional methods and deep learning methods.
[0004] Traditional methods mainly include generative and discriminative methods. Generative methods are committed to building a target appearance model and finding the area that best matches the model in the subsequent frames of the video sequence. Typical algorithms include optical flow, Kalman filtering, and Meanshift, which have made significant progress in the field of target tracking. However, generative methods rely too much on the target itself in feature extraction and ignore background information, which leads to template drift or target loss when the target appearance changes dramatically or is occluded, limiting its application scenarios. Discriminative methods, especially correlation filtering algorithms, maximize the correlation between the target area and the search area by training filters, and then locate the area with the strongest response in the next frame as the target position. Since the MOSSE algorithm was proposed in 2010, the correlation filtering algorithm has attracted widespread attention for its high efficiency and good scene adaptability. Subsequently, the CSK algorithm proposed by Henriques et al. introduced dense sampling and kernelized correlation filtering based on MOSSE, and the KCF algorithm incorporated the HOG (Histogram of Oriented Gradients) feature of multi-channel gradients. In order to continuously improve the performance of the correlation filtering algorithm, researchers have introduced methods such as multi-feature fusion and scale adaptation. For example, the SAMF algorithm and the DSST algorithm respectively introduce scale factors and scale filters to enhance the adaptability to the change of target scale; the Staple algorithm combines color features on the basis of DSST and uses color histograms to improve the tracking ability of deformed targets. In addition, some innovative methods, such as the MEEM algorithm, improve the tracking performance by optimizing edge metrics; the Struck algorithm uses support vector machines for target tracking; the TLD algorithm integrates tracking, learning and detection strategies to maintain the robustness of long-term tracking; the CACF algorithm adds background information in four directions to the training to improve the tracking accuracy of the target under rapid movement. Despite this, traditional methods still have limitations when dealing with complex scenes. For example, the target reappears after a short disappearance, the target undergoes drastic deformation, and the tracking effect is not satisfactory in scenes such as rapid movement.
[0005] In recent years, deep learning-based target tracking methods have also made significant progress. Compared with traditional methods, the powerful feature extraction capabilities of deep learning methods can automatically learn richer feature representations and better cope with changes in target appearance. Some representative deep learning tracking algorithms, such as CNN-based CFNet and SiamFC, CNN-Transformer-based BANDT, and Transformer-based SwinTrack, have demonstrated excellent performance on multiple standard datasets. However, deep learning methods still face challenges in data generalization and robustness, especially when the amount of data is limited or the application scenario changes.
[0006] The Staple algorithm combines two different tracking models, the correlation filter and the color histogram model, to improve the accuracy and robustness of tracking. The Staple algorithm also inherits the scale adaptation of the DSST algorithm and better adapts to changes in target size. However, in actual scenarios, there are several problems:
[0007] 1. When the target is significantly deformed and the background and target are similar in color, the background and target areas cannot be accurately distinguished, which reduces the tracking performance;
[0008] 2. In the Staple algorithm framework, the maximum value of the combined response of the HOG feature and the color histogram is usually located as the target center. However, in practical applications, when multiple maxima appear in the mixed response, or there are other targets with similar appearance to the target, a single maximum criterion may not be able to accurately identify the true target center.
[0009] 3. How to judge the target loss at the right time is a major challenge in the field of visual tracking. Common evaluation indicators include maximum response value, peak sidelobe ratio, average peak correlation energy and spatial reliability. However, the Staple algorithm does not have a loss judgment mechanism. When the target is lost or blocked, the tracker of the Staple algorithm will continue to update the template, which may lead to error accumulation and significantly reduce the probability of successful detection when the target reappears.
[0010] 4. In various application scenarios, HOG features and color features have their own advantages in terms of credibility when describing the target state. For example, when the target is deformed, color features may provide a more accurate description; on the contrary, when the background color interference is more serious, HOG features may be more accurate and reliable. Therefore, when the scene changes, if only a single feature description method is used, it will inevitably lead to a decrease in tracking accuracy. Summary of the invention
[0011] The present invention aims to solve the above technical problems in the prior art and provide a complex scene single target tracking method, device, system and storage medium based on the Staple algorithm. The tracking method of the present invention is a novel complex scene single target tracking method based on the Staple algorithm, which is used to improve the target tracking performance in complex scenes such as target deformation, occlusion, similar interference and field of view exceeding.
[0012] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0013] A complex scene single target tracking method based on Staple algorithm includes the following steps:
[0014] Step 1: Read a new frame image and determine whether the current frame is the starting frame. If so, select the tracking target and initialize the HOG feature model, scale-related filter model, foreground histogram, and background weight histogram model, and return to Step 1; if not, proceed directly to Step 2;
[0015] Step 2: Calculate the HOG feature response of this frame using the target center position obtained in the previous frame, the HOG feature model and the scale-related filter model; Calculate the foreground probability of this frame using the foreground histogram and background weight histogram model obtained in the previous frame;
[0016] Step 3: Perform feature adaptive fusion on the HOG feature response and foreground probability of the frame obtained in Step 2;
[0017] Step 4: Re-identify similar targets based on the mixed features fused in Step 3, eliminate the interference of similar targets, and determine the true target center position;
[0018] Step 5: Use the loss determination mechanism module to determine whether the target is lost, ensuring that the relevant model can accurately identify the real target rather than continuously updating with the wrong target; if the target is lost, return to Step 1; otherwise, proceed to Step 6;
[0019] Step 6: Estimate the scale of the target detected in Step 5 to obtain a bounding box that matches the target size;
[0020] Step 7: Based on the current frame detection results obtained in the previous steps, the template update strategy is used to update the HOG feature model, scale-related filter model, and background weight histogram model according to fixed weights;
[0021] Step 8: If the current frame is not the last frame, return to Step 1 until the current frame is the last frame image and exit the program.
[0022] In the above technical solution, the background weight histogram model calculation in Step 2 uses a two-dimensional Gaussian function to set the background pixel weight according to the distance between the background area pixel and the target center point. The steps are as follows:
[0023] Step a: Take the target center position detected in the previous frame as the origin and the length and width of the search area as the size to establish a two-dimensional Gaussian function as follows:
[0024]
[0025] Where W x,y is the weight of the pixel value at (x, y), (x0, y0) is the center position of the target detected in the previous frame, (x, y) is the background pixel position, δw and δ h is the standard deviation of the two-dimensional Gaussian function, which is calculated as follows:
[0026]
[0027] w and h are the length and width of the search area respectively, α w , α h They are used to limit δ w and δ h The coefficient of size is used to reasonably adjust the weight of the edge pixels of the search area according to the moving speed of the target in actual application;
[0028] Step b: Normalize the background area weight obtained, the formula is as follows:
[0029]
[0030] In the formula, is the normalized weight of the (x, y) position;
[0031] Step c: Calculate the background area weight histogram H(b), the formula is as follows:
[0032]
[0033] Where b is the number of histogram channels corresponding to different color values, Q(·) represents the statistical color histogram, and x x,y Represents the pixel value at the (x, y) position.
[0034] In the above technical solution, the feature adaptive fusion in Step 3 uses the peak sidelobe ratio PSR to quantitatively evaluate the reliability of the HOG feature response and the color feature response to obtain the fusion weight, and through linear fusion, obtains the adaptively fused mixed feature response Mix_res.
[0035] In the above technical solution, the adaptively fused mixed feature response Mix_res obtained in the above steps has the following fusion formula:
[0036] Mix_res=(1-σ)×cf+σ×pwp(9)
[0037] In the formula, cf is the HOG feature response, pwp is the color feature response, and σ is the fusion coefficient;
[0038] By calculating the PSR values of cf and pwp, the credibility of different features can be determined, and the fusion coefficient σ is constructed accordingly. The calculation method is shown in formula (10):
[0039]
[0040] Where PSR_cf is the PSR value of the HOG feature response cf, PSR_pwp is the PSR value of the color feature response pwp, and the PSR value is an indicator to measure the strength relationship between the maximum response value and the peak sidelobe ratio. The larger the value, the greater the probability that the point is the target. The calculation formula is as follows:
[0041]
[0042] In the formula, cf max is the maximum value of the HOG feature response, R cf It is the HOG feature response cf in cf max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R cf ) is R cf The standard deviation of the response values in the set; pwp max is the maximum value of the color feature response pwp, R pwp Is the color feature response pwp in pwp max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R pwp ) is used to calculate R pwp The standard deviation of the response values in the set.
[0043] In the above technical solution, the specific steps of similar target re-identification in Step 4 are as follows:
[0044] Step a: Extract the local maximum max L from the mixed feature response map i , i = 1, ..., I, I represents the number of local maxima, and calculates these local maxima max L i , i=1,..., the ratio of I to the global maximum value max G;
[0045] Step b: Determine whether similar target re-identification is needed;
[0046] By setting a threshold β, if the ratio of the local maximum to the global maximum does not exceed this threshold, the position corresponding to the global maximum is considered to be the target center position, and similar targets are no longer re-identified; if the ratio of the local maximum to the global maximum exceeds the threshold, it is considered that there is similar target interference, and similar targets are re-identified using step c;
[0047] Step c: Select the local maximum max l filtered by the threshold p The corresponding position is the center, p = 1, ..., P, P represents the number of local maxima after screening, and the area is η times the size of the target area. The characteristic response and global maximum value max g in this area are calculated. p and its position (x p ,yp );
[0048] Step d: To ensure the consistency of response intensity and target space, a weighted fusion strategy of target center distance and maximum response value is designed to obtain the final response value score p , as shown in formula (5):
[0049]
[0050] In the formula, λ is the weight factor of weighted fusion, d p is the local response maximum value max l p Corresponding position (x p ,y p ) and the Euclidean distance between the target center position (x0, y0) in the previous frame, which is calculated as shown in formula (6):
[0051]
[0052] Step e: Get score p , p = 1, ..., the maximum value of P max g p The corresponding position (x p ,y p ) is the center position of the tracking target.
[0053] In the above technical solution, the loss judgment mechanism module in Step 5 uses the maximum value of the mixed feature response maxG, the maximum value of the HOG feature response max CF and the average peak value apce of the correlation energy of the mixed response as judgment criteria.
[0054] In the above technical solution, the loss determination mechanism module in the above Step 5 determines whether the target is lost. The specific determination method is as follows:
[0055]
[0056] Where frame is the current frame, frames is the frame number from the first frame to accurately track the target, μ1, μ2, μ3 are weight coefficients set according to the actual application scenario, and apce is the average peak correlation energy of the M×N regional response map, which is calculated as follows:
[0057]
[0058] Where Mix_res is the mixed response diagram, Mix_res m,n is the mixed response value of the pixel (m, n).
[0059] In the above technical solution, the template update strategy in Step 7 is based on the loss judgment, and the maximum value of the mixed feature response Mix_res and the average peak energy apce are added as the update basis.
[0060] In the above technical solution, the specific method of the template update strategy in the above Step 7 is shown in formula (12):
[0061]
[0062] In the formula, frame is the current frame, frames is the frame number from the first frame to accurately track the target. It is a weight coefficient set according to the actual application scenario. When A∧(BV C) is satisfied, it is considered that the tracking effect of the frame is relatively good, and the template is updated. The update strategy of the Staple algorithm is adopted during the update, in which the HOG feature response and the color feature response are updated according to specific learning rates respectively.
[0063] A device, system or storage medium for running the complex scene single target tracking method based on the Staple algorithm of the present invention.
[0064] The beneficial effects of the present invention are:
[0065] The complex scene single target tracking method based on Staple algorithm of the present invention makes innovations in background weight histogram, similar target re-identification, loss judgment, feature adaptive fusion and template updating. The application of background weight histogram, on the one hand, enhances the suppression of background pixels in the target background area by assigning higher weights to pixels closer to the target center; on the other hand, by assigning lower weights to pixels far from the target center, the interference of the background area far from the target is weakened. This weight distribution mechanism effectively enhances the contrast between the target and the background. The application of similar target re-identification takes into account the influence of two factors: the response intensity of the target and the motion coherence of the target in the video sequence. The target center distance and the maximum mixed response value between the two frames are used to eliminate similar targets, ensuring that the tracking algorithm focuses on the correct target. The loss judgment mechanism effectively avoids the accumulation of errors of the tracker during target loss or occlusion, and can significantly improve the probability of successful detection and reconstruction of the target after loss or occlusion. Feature adaptive fusion dynamically adjusts the fusion ratio of HOG features and color features based on the peak sidelobe ratio PSR to adapt to feature extraction under target deformation or background color interference, which is conducive to tracking accuracy. The template update strategy reasonably selects the template update timing based on the mixed response and the average peak energy APCE to achieve more accurate target state description and tracking. Through the integration and improvement of these innovations, the present invention not only improves the accuracy of the algorithm, but also significantly enhances its adaptability and robustness in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 This is a flowchart of the single target tracking method in complex scenes based on the Staple algorithm.
[0067] Figure 2 The success rate and accuracy curves of 14 algorithms on the selected data set, where (a) is the success rate curve and (b) is the accuracy curve.
[0068] Figure 3 The success rate and accuracy curves based on deformation attributes, where (a) is the success rate curve and (b) is the accuracy curve.
[0069] Figure 4 The success rate and accuracy curves based on occlusion attributes, where (a) is the success rate curve and (b) is the accuracy curve.
[0070] Figure 5 The success rate and accuracy curves based on the out-of-view attribute, where (a) is the success rate curve and (b) is the accuracy curve.
[0071] Figure 6 The following is a comparison chart of the effects of different algorithms on the Bird1 video sequence.
[0072] Figure 7 The following is a comparison chart of the effects of different algorithms on the Girl2 video sequence.
[0073] Figure 8 The following is a comparison chart of the effects of different algorithms on the Tiger2 video sequence. DETAILED DESCRIPTION
[0074] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0075] The complex scene single target tracking method based on the Staple algorithm of the present invention optimizes the three key links of weighted histogram, similar target feature re-identification and loss judgment, and is supplemented by feature adaptive fusion and template updating strategy, aiming to achieve accurate tracking of targets in complex scenes.
[0076] The complex scene single target tracking method based on Staple algorithm of the present invention is described in detail in Figure 1 , the specific steps are as follows:
[0077] Step 1: Read a new frame image and determine whether the current frame is the starting frame. If so, select the tracking target and initialize its HOG feature model, scale-related filter model, foreground histogram, and background weight histogram model, and then return to (1); if not, proceed directly to Step 2.
[0078] Step 2: Calculate the HOG feature response of this frame using the target center position, size and HOG feature filter template obtained in the previous frame; Calculate the foreground probability of this frame using the foreground and background weight histogram (1) obtained in the previous frame.
[0079] Step 3: Perform feature adaptive fusion (2) on the HOG feature response and foreground probability of the frame.
[0080] Step 4: Re-identify similar targets (3) based on the mixed features fused in Step 3 to eliminate the interference of similar targets and determine the true target center position.
[0081] Step 5: Use the loss determination mechanism module (4) to determine whether the target is lost, ensuring that the model can accurately identify the true target rather than continuously updating with the wrong target. If the target is lost, return to Step 1; otherwise, proceed to Step 6.
[0082] Step 6: Estimate the scale of the detected target to obtain a bounding box that matches the target size.
[0083] Step 7: Based on the detection result of the current frame, the three models (HOG features, scale-related filter model and background weight histogram) are updated using the template update strategy (5) according to the fixed weights.
[0084] Step 8: If the current frame is not the last frame, return to Step 1 until the current frame is the last frame image and exit the program.
[0085] In the tracking method of the present invention, the background weight histogram calculation in Step 2 is specifically as follows:
[0086] In order to more accurately distinguish the background from the target, especially when the target is significantly deformed and the background and target are similar in color, the farther the pixel is from the center of the target, the less likely it is to be the target. Therefore, the present invention uses a two-dimensional Gaussian function to set the background pixel weight according to the distance between the background area pixel and the target center point, so as to more accurately identify and distinguish the background from the target, thereby improving the accuracy and robustness of the tracking algorithm. The specific steps are as follows:
[0087] Step 1: Take the target center position detected in the previous frame as the origin and the length and width of the search area as the size to establish a two-dimensional Gaussian function:
[0088]
[0089] Where W x,yis the weight of the pixel value at (x, y), (x0, y0) is the target center detected in the previous frame, (x, y) is the background pixel position, δ w and δ h is the standard deviation of the two-dimensional Gaussian function, which is calculated as follows:
[0090]
[0091] Where w and h are the length and width of the search area, respectively, and α w , α h They are used to limit δ w and δ h The coefficient of size is used to reasonably adjust the weight of the edge pixels in the search area according to the moving speed of the target in actual application.
[0092] Step 2: Normalize the background area weights obtained:
[0093]
[0094] In the formula, is the weight after normalization of the (x, y) position. Weight normalization is used to balance the weight contribution of different regions, ensuring that when performing target and background segmentation, the model can focus more on the target area while suppressing the influence of background noise.
[0095] Step 3: Calculate the background area weight histogram H(b):
[0096]
[0097] Where b is the number of histogram channels corresponding to different color values, Q(·) represents the statistical color histogram, and z x,y Represents the pixel value at the (x, y) position.
[0098] In the tracking method of the present invention, the feature adaptive fusion process in Step 3 is specifically as follows:
[0099] In various application scenarios, HOG features and color features have their own advantages in terms of credibility when describing the target state. For example, when the target is deformed, the color feature may provide a more accurate description; on the contrary, when the background color interference is more serious, the HOG feature may be more accurate and reliable. In order to improve the accuracy of feature response, the present invention designs an adaptive weight strategy, that is, dynamically adjusts the weight of the feature according to its reliability, so as to enhance the influence of more reliable features.
[0100] The present invention uses the peak sidelobe ratio PSR to quantitatively evaluate the reliability of the HOG feature response and the color feature response and then obtains the fusion weight. Through linear fusion, the adaptive fused mixed feature response Mix_res is obtained. The fusion formula is as follows:
[0101] Mix_res=(1-σ)×cf+σ×pwp(9)
[0102] In the formula, cf is the HOG feature response, pwp is the color feature response, and σ is the fusion coefficient. By calculating the PSR values of cf and pwp, the credibility of different features can be determined, and the fusion coefficient σ is constructed accordingly. The calculation method is shown in formula (10):
[0103]
[0104] Where PSR_cf is the PSR value of the HOG feature response cf, PSR_pwp is the PSR value of the color feature response pwp, and the PSR value is an indicator to measure the strength relationship between the maximum response value and the peak sidelobe ratio. The larger the value, the greater the probability that the point is the target. The calculation formula is as follows:
[0105]
[0106] In the formula, cf max is the maximum value of the HOG feature response, R cf It is the HOG feature response cf in cf max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R cf ) is R cf The standard deviation of the response values in the set; pwp max is the maximum value of the color feature, R pwp Is the color feature response pwp in pwp max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R pwp ) is used to calculate R pwp The standard deviation of the response values in the set.
[0107] In the tracking method of the present invention, the specific process of similar target re-identification in Step 4 is as follows:
[0108] In the Staple algorithm framework, the maximum value of the combined response of the HOG feature and the color histogram is usually positioned as the target center. However, in practical applications, when multiple maxima appear in a mixed response, or when there are other targets with similar appearance to the target, a single maximum criterion may not be able to accurately identify the true target center. In view of the above two points, the present invention designs a similar target re-identification module, which aims to enhance the recognition ability of the true target. When there are multiple peaks in the mixed response, the module further analyzes the pattern of the feature response to assist in determining which peak is more likely to be the true target response. The specific steps are as follows:
[0109] Step 1: Extract the local maximum max L from the mixed feature response map i , i = 1, ..., I (I represents the number of local maxima), and calculate these local maxima max L i , i=1,..., the ratio of I to the global maximum value max G.
[0110] Step 2: Determine whether similar target re-identification is needed. By setting a threshold β, if the ratio of the local maximum to the global maximum does not exceed this threshold, the position corresponding to the global maximum is considered to be the target center position, and similar target re-identification is no longer performed; if it exceeds the threshold, it is considered that there is similar target interference, and Step 3 is used to re-identify similar targets.
[0111] Step 3: Select the local maximum value max l filtered by the threshold p , p=1,...,P (P represents the number of local maxima after screening) corresponds to the position centered at n times the target area, and calculates the characteristic response and global maximum value max g in the area p and its position (x p ,y p ).
[0112] Step 4: To ensure the consistency of response intensity and target space, a weighted fusion strategy of target center distance and maximum response value is designed to obtain the final response value score p As shown in formula (5):
[0113]
[0114] In the formula, λ is the weight factor of weighted fusion, d p is the local response maximum value max l p Corresponding position (x p ,y p ) and the Euclidean distance between the target center position (x0, y0) in the previous frame, which is calculated as shown in formula (6):
[0115]
[0116] Step 5: Get score p , p = 1, ..., the maximum value of P max g p The corresponding position (x p ,y p ) is the center position of the tracking target.
[0117] In the tracking method of the present invention, the specific evaluation method of the loss determination mechanism in Step 5 is as follows:
[0118] How to judge the loss of the target at the right time is a major challenge in the field of visual tracking. Commonly used evaluation indicators include maximum response value, peak-to-sidelobe ratio, average peak correlation energy and spatial reliability. However, the Staple algorithm does not have a loss judgment mechanism. When the target is lost or occluded, the tracker of the Staple algorithm will continue to update the template, which may lead to error accumulation and significantly reduce the probability of successful detection of the target when it reappears. To solve this problem, the present invention designs a new loss judgment mechanism. The mechanism is based on the continuity between frames and uses the information of 5 consecutive frames of images as the basis for judging the loss of the target, so as to more accurately capture the change of the target state. In addition, in order to comprehensively evaluate the loss or occlusion of the target, the maximum value of the mixed response maxG, the maximum value of the HOG feature response maxCF and the average peak value of the correlation energy APCE (Average Peak-to-Correlation Energy) of the mixed response are used as the judgment criteria (the average peak value of the correlation energy APCE is indicated by apce below). The specific judgment method is as follows:
[0119]
[0120] In the formula, frame is the current frame, frames is the frame number from the first frame to accurately track the target, μ1, μ2, μ3 are weight coefficients set according to the actual application scenario, and apce is the average peak correlation energy of the response map of the M×N area, which represents the fluctuation degree of the response map. The larger the value, the greater the fluctuation, and the greater the probability that the point is the target. The calculation formula is as follows:
[0121]
[0122] Where Mix_res is the mixed feature response map, Mix_res m,n is the mixed response value of the pixel (m, n).
[0123] When the target is judged to be lost or blocked according to the above evaluation method, the strategy adopted is to keep the position and size of the target box consistent with the previous frame until the target is successfully detected again. At this time, the update operation of the tracker template is suspended to exclude any irrelevant information that may interfere with the tracking accuracy. This measure is intended to maintain the stability of the tracker performance, increase the probability of accurate detection when the target reappears, and ensure the continuity and reliability of the entire tracking process.
[0124] In the tracking method of the present invention, the specific method of the template update strategy in Step 7 is as follows:
[0125] The performance of the correlation filter tracking algorithm depends largely on the template accuracy, that is, the current state of the target model. When the target moves quickly or undergoes significant deformation, if the template update strategy is not flexible enough, the tracking performance will be degraded; or when the target is lost or occluded, if the template is still updated, it will cause the template to drift. To solve this problem, the present invention improves the template update strategy of the Staple algorithm to enhance the adaptability and robustness of the algorithm.
[0126] The present invention adds the maximum value of the mixed characteristic response Mix_res and the average peak energy apce as the update basis based on the loss judgment. The specific method is shown in formula (12):
[0127]
[0128] In the formula, frame is the current frame, frames is the frame number from the first frame to accurately track the target. It is a weight coefficient set according to the actual application scenario. When A∧(BVC) is satisfied, it is considered that the tracking effect of the frame is good and the template can be updated. The update strategy of the Staple algorithm is adopted during the update, in which the HOG feature is updated at a learning rate of 0.01 and the color feature is updated at a learning rate of 0.04.
[0129] In order to better judge the performance of the complex scene single target tracking method based on the Staple algorithm of the present invention, the experimental data set uses 64 video sequences (OTB64) in OTB100
[23] , which mainly include attributes such as occlusion, deformation, out of field of view, and interference from similar targets. A comprehensive comparative analysis is carried out with 13 classic tracking algorithms such as Staple, CFNet, and KCF from both quantitative and qualitative aspects.
[0130] (1) Quantitative analysis
[0131] The success rate and precision are used as quantitative analysis indicators of the tracking algorithm, where the success rate (SR) refers to the proportion of frames in which the IOU between the target real frame and the target tracking frame is greater than a given threshold; the precision (P) refers to the proportion of frames in which the difference between the target center position tracked by the algorithm and the actual target center position is less than a certain threshold. The One-Pass Evaluation (OPE) method is used to experimentally evaluate the relevant tracking algorithms. The success rate and precision of the method proposed in this invention are compared with the selected dataset as a whole and under the three attributes of deformation, occlusion, and out of view. The results are shown in Tables 1 and 2. Figure 2-5 Displayed in.
[0132] Table 1 Tracking results of different algorithms on OTB64
[0133]
[0134] As shown in Table 1, the method proposed in the present invention ("Our work" is used in the figures and tables to represent the method of the present invention) has improved success rates compared with the Staple algorithm on 64 videos involving the three attributes of deformation, occlusion, and out of view in the OTB100 dataset, with specific improvements of 1.6%, 3.3%, 1%, and 5.5%, respectively. In addition, compared with some deep learning-based methods such as SiamFC and CFNet, the success rate of the method proposed in the present invention has also been improved to a certain extent. Among all the algorithms involved in the comparison, the success rates of the method proposed in the present invention ranked first, first, and second in the attributes of deformation, occlusion, and out of view, respectively, and the overall performance ranked first; in terms of accuracy, the method proposed in the present invention improved the deformation and out of view attributes by 2.8% and 1.2% respectively compared with the Staple algorithm.
[0135] (2) Qualitative analysis
[0136] In order to more intuitively demonstrate the effect comparison between the method proposed in this invention and other methods, a series of challenging video sequences are selected for analysis. These videos are representative in terms of target deformation, occlusion, and out-of-field properties, which can fully test the robustness and adaptability of the tracking algorithm. Figure 6 The Bird1 video represents the deformation attribute. Figure 7 The Girl2 video in the middle represents the occlusion attribute. Figure 8 The Tiger2 video in the figure represents the target out of view attribute. In all the schematic diagrams, the purple box is the true value, the red box is the algorithm of the present invention, the green box is the Staple algorithm, the blue box is the MEEM algorithm, and the black box is the SiamFC algorithm.
[0137] Figure 6The bird in the middle of the 214th frame of the Bird1 video is constantly changing shape, and in subsequent frames, interference objects similar to the target appear in the picture. When processing this complex scene, the algorithm of the present invention can continue to stably follow the movement of the target. Although there is still a certain gap from the true value, its effect is significantly enhanced compared with other algorithms. This scene verifies the anti-deformation and anti-interference properties of the algorithm of the present invention. It can well adapt to changes in the appearance of the target and resist interference in similar backgrounds.
[0138] Figure 7 The target girl in the Girl2 video is shown to be blocked by a passing man in the 105th and subsequent frames of the video, causing the target to be temporarily invisible in the video. In the 105th frame, all algorithms can accurately frame the target, but with the blocking of the man, all algorithms except the proposed algorithm will track the man, and will not return to the target until the girl appears in the detection frame again. This scene verifies the effectiveness of the algorithm proposed in the present invention for occlusion. Compared with other algorithms, the method of the present invention shows higher tracking accuracy and robustness when the target is blocked and reappears.
[0139] Figure 8 It shows that the ragdoll tiger appears partially out of the field of view in the 86th frame. When the target is still in the field, the tracking effects of all algorithms are not much different, but in the 86th frame, it can be seen that the algorithm proposed by the present invention overreacts to the action of the target exceeding the field of view, and the target frame is almost out of the field of view, but fortunately, when the target reappears in the 92nd frame, it can adapt quickly and maintain tracking compared with other algorithms. Although the response of the tracking frame is slightly excessive when the target is partially out of the field of view, the algorithm can take advantage of its design and quickly and accurately resume tracking after the target re-enters the field of view. This scene verifies that the algorithm proposed by the present invention has good adaptability and recovery capabilities for targets that reappear beyond the field of view.
[0140] The complex scene single target tracking method based on the Staple algorithm proposed in the present invention can be applied not only in embedded chips, but also embedded in host computer software, such as a device, equipment or storage medium for running the tracking method of the present invention.
[0141] In summary, the complex scene single target tracking method based on the Staple algorithm of the present invention makes innovations in background weight histogram, similar target re-identification, loss judgment, feature adaptive fusion and template updating.
[0142] (1) Background weight histogram: By assigning higher weights to pixels closer to the target center, the suppression of background pixels in the target background area is enhanced; by assigning lower weights to pixels farther from the target center, the interference of background areas farther from the target is weakened. This weight distribution mechanism effectively enhances the contrast between the target and the background.
[0143] (2) Adaptive feature fusion: The fusion ratio of HOG features and color features is dynamically adjusted based on the peak sidelobe ratio (PSR). This is to adapt to feature extraction when the target is deformed or the background color is disturbed, which is beneficial to the tracking accuracy.
[0144] (3) Similar target re-identification: Taking into account the influence of both the target’s response strength and the target’s motion coherence in the video sequence, similar targets are eliminated using the target center distance and the maximum mixed response value between two frames to ensure that the tracking algorithm focuses on the correct target.
[0145] (4) Loss judgment mechanism: (a) Based on the continuity between frames, the information of 5 consecutive frames of images is used as the basis for judging target loss, so as to more accurately capture the target state changes; (b) The maximum value of the hybrid feature response max G, the maximum value of the HOG feature response max CF and the average peak value of the correlation energy APCE (Average Peak-to-Correlation Energy) of the hybrid feature response are used as judgment criteria to judge the loss or occlusion.
[0146] When the target is judged to be lost or occluded according to the above evaluation method, the strategy adopted is to keep the position and size of the target box consistent with the previous frame until the target is successfully detected again. At this time, the update operation of the tracker template is suspended to exclude any irrelevant information that may interfere with the tracking accuracy. This measure is intended to maintain the stability of the tracker performance, increase the probability of accurate detection when the target reappears, and ensure the continuity and reliability of the entire tracking process. It effectively avoids the accumulation of errors during the period when the tracker is lost or occluded, and can significantly increase the probability of successful detection and reconstruction of the target after loss or occlusion.
[0147] (5) The template update strategy reasonably selects the template update timing based on the mixed feature response and average peak energy APCE to achieve more accurate target state description and tracking.
[0148] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.
Claims
1. A complex scene single target tracking method based on Staple algorithm, characterized in that: The following steps are involved: Step 1: Read a new frame image and determine whether the current frame is the starting frame. If so, select the tracking target and initialize the HOG feature model, scale-related filter model, foreground histogram, and background weight histogram model, and return to Step 1; if not, proceed directly to Step 2; Step 2: Calculate the HOG feature response of this frame using the target center position obtained in the previous frame, the HOG feature model and the scale-related filter model; Calculate the foreground probability of this frame using the foreground histogram and background weight histogram model obtained in the previous frame; Step 3: Perform feature adaptive fusion on the HOG feature response and foreground probability of the frame obtained in Step 2; Among them, the feature adaptive fusion uses the peak sidelobe ratio PSR to quantitatively evaluate the reliability of the HOG feature response and the color feature response, and then obtains the fusion weight. Through linear fusion, the adaptive fusion mixed feature response Mix_res is obtained; The adaptive fusion mixed feature response Mix_res has the following fusion formula: Mix_res=(1-σ)×cf+σ×pwp (9) In the formula, cf is the HOG feature response, pwp is the color feature response, and σ is the fusion coefficient; By calculating the PSR values of cf and pwp, the credibility of different features can be determined, and the fusion coefficient σ is constructed accordingly. The calculation method is shown in formula (10): Where PSR_cf is the PSR value of the HOG feature response cf, PSR_pwp is the PSR value of the color feature response pwp, and the PSR value is an indicator to measure the strength relationship between the maximum response value and the peak sidelobe ratio. The larger the value, the greater the probability that the point is the target. The calculation formula is as follows: In the formula, cf max is the maximum value of the HOG feature response, R cf It is the HOG feature response cf in cf max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R cf ) is R cf The standard deviation of the response values in the set; pwp max is the maximum value of the color feature response pwp, R pwp Is the color feature response pwp in pwp max The corresponding pixel is the set of response values in the area outside the 11×11 neighborhood of the center, std(R pwp ) is used to calculate R pwp The standard deviation of the response values in the set; Step 4: Re-identify similar targets based on the mixed features fused in Step 3, eliminate the interference of similar targets, and determine the true target center position; Step 5: Use the loss determination mechanism module to determine whether the target is lost, ensuring that the relevant model can accurately identify the real target rather than continuously updating with the wrong target; if the target is lost, return to Step 1; otherwise, proceed to Step 6; Among them, the loss judgment mechanism module uses the maximum value of the mixed feature response max G, the maximum value of the HOG feature response max CF and the average peak value apce of the correlation energy of the mixed response as the judgment criteria; Step 6: Estimate the scale of the target detected in Step 5 to obtain a bounding box that matches the target size; Step 7: Based on the current frame detection results obtained in the previous steps, the template update strategy is used to update the HOG feature model, scale-related filter model, and background weight histogram model according to fixed weights; Among them, the template update strategy is based on the loss judgment, and the maximum value of the mixed feature response Mix_res0 and the average peak energy apce are added as the update basis; Step 8: If the current frame is not the last frame, return to Step 1 until the current frame is the last frame image and exit the program.
2. The complex scene single target tracking method based on Staple algorithm according to claim 1 is characterized in that: In Step 2, the background weight histogram model is calculated using a two-dimensional Gaussian function, and the background pixel weight is set according to the distance between the background area pixel and the target center point; the steps are as follows: Step a: Take the target center position detected in the previous frame as the origin and the length and width of the search area as the size to establish a two-dimensional Gaussian function as follows: Where W x,y is the weight of the pixel value at (x, y), (x0, y0) is the center position of the target detected in the previous frame, (x, y) is the background pixel position, δ w and δ h is the standard deviation of the two-dimensional Gaussian function, which is calculated as follows: w and h are the length and width of the search area respectively, α w , α h They are used to limit δ w and δ h The coefficient of size is used to reasonably adjust the weight of the edge pixels of the search area according to the moving speed of the target in actual application; Step b: Normalize the background area weight obtained, the formula is as follows: In the formula, is the normalized weight of the (x,y) position; Step c: Calculate the background area weight histogram H(b), the formula is as follows: Where b is the number of histogram channels corresponding to different color values, Q(·) represents the statistical color histogram, and z x,y Represents the pixel value at the (x,y) position.
3. The complex scene single target tracking method based on Staple algorithm according to claim 1, characterized in that: The specific steps of similar target re-identification in Step 4 are as follows: Step a: Extract the local maximum max L from the mixed feature response map i , i=1,…,I, I represents the number of local maxima, and calculates these local maxima max L i ,i=1,…,the ratio of I to the global maximum value max G; Step b: Determine whether similar target re-identification is needed; By setting a threshold β, if the ratio of the local maximum to the global maximum does not exceed this threshold, the position corresponding to the global maximum is considered to be the target center position, and similar targets are no longer re-identified; if the ratio of the local maximum to the global maximum exceeds the threshold, it is considered that there is similar target interference, and similar targets are re-identified using step c; Step c: Select the local maximum max l filtered by the threshold p The corresponding position is the center, p = 1, ..., P, P represents the number of local maxima after screening, and the area is η times the size of the target area. The characteristic response and global maximum value max g in this area are calculated. p and its position (x p ,y p ); Step d: To ensure the consistency of response intensity and target space, a weighted fusion strategy of target center distance and maximum response value is designed to obtain the final response value score p , as shown in formula (5): In the formula, λ is the weight factor of weighted fusion, d p is the local response maximum value max l p Corresponding position (x p ,y p ) and the Euclidean distance between the target center position (y0,y0) in the previous frame, which is calculated as shown in formula (6): Step e: Get score p ,p=1,…,the maximum value of P max g p The corresponding position (x p ,y p ) is the center position of the tracking target.
4. The complex scene single target tracking method based on Staple algorithm according to claim 1 is characterized in that: The specific evaluation method in Step 5 is as follows: Where frame is the current frame, frames is the frame number from the first frame to accurately track the target, μ1, μ2, μ3 are weight coefficients set according to the actual application scenario, and apce is the average peak correlation energy of the M×N regional response map, which is calculated as follows: Where Mix_res is the mixed response diagram, Mix_res m,n is the mixed response value of the pixel (m,n).
5. The complex scene single target tracking method based on Staple algorithm according to claim 1, characterized in that: The specific method of the template update strategy in Step 7 is shown in formula (12): In the formula, frame is the current frame, frames is the frame number from the first frame to accurately track the target. It is a weight coefficient set according to the actual application scenario. When A∧(B∨C) is satisfied, the tracking effect of the frame is considered good and the template is updated. The update strategy of the Staple algorithm is adopted during the update, in which the HOG feature response and the color feature response are updated according to specific learning rates respectively.
6. A device, system or storage medium for running the complex scene single target tracking method based on the Staple algorithm as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Long-term target tracking method and system based on improved Staple
CN115049706A
Anti-occlusion target tracking algorithm based on correlation filtering Staple algorithm
CN117934548A