A Robust Visual Object Tracking Algorithm from Coarse to Fine Based on Sparse Learning

By introducing sparse learning and interactive Kalman filtering methods in visual target tracking, problems such as target and background information processing, feature diversity and redundancy, and response positioning accuracy in the prior art are solved, and more efficient and robust visual target tracking is achieved.

CN115294174BActive Publication Date: 2025-06-27CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211011081.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-06-27
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

The existing visual target tracking methods have shortcomings in dealing with problems such as target and surrounding background information, target feature diversity and redundancy, maximum positioning accuracy of target responses, and model drift, resulting in poor tracking performance.

Method used

A robust visual target tracking algorithm based on sparse learning is adopted, by introducing regular constraints on the basis of classical filtering, using lasso regression to select features with forward reference value, the sparsity of the target response is embedded into the tracking framework, and interactive Kalman filtering is used for supervision and positioning.

Benefits of technology

It effectively suppresses interference such as target scale changes, background clutter, and lighting changes, avoids model drift and tracking failures, and improves the accuracy and robustness of target positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294174B_ABST
    Figure CN115294174B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of visual target tracking, and specifically provides a robust visual target tracking algorithm from coarse to fine based on sparse learning, including step S1: establishing a target appearance model, and S2: tracking the target. In step S1, context constraint and lasso regression are introduced, and the sparsity of the target response is embedded into the tracking framework; in step S2, a coarse-to-fine positioning method is adopted, and interactive Kalman filtering is used to supervise and participate in the positioning of the target, and multi-peak detection is performed on the response map during fine positioning. The target tracking algorithm provided by the present invention can effectively suppress the influence of internal and external interferences of the target by establishing a robust and sparse target appearance model, effectively avoid the occurrence of model drift or tracking failure caused by accidental peaks, and can improve the accuracy of positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual tracking, and particularly to a robust visual object tracking algorithm from coarse to fine based on sparse learning. Background Art

[0002] Visual object tracking is a basic research topic in the field of computer vision, which is widely applied not only in civilian fields such as intelligent video surveillance and human-computer interaction, but also in military fields such as aerial reconnaissance and precision guidance. Its goal is to give the state of any object in the initial frame of a video sequence and continuously locate the state of the object in subsequent video sequences.

[0003] Visual object tracking methods can be roughly divided into two categories: generative and discriminative, according to the way of establishing the object appearance model. The generative method first models the object appearance, and then locates the object by maximizing the similarity between the object candidate region and the object model or minimizing the corresponding reconstruction error. For a long time, existing visual object tracking mostly uses the generative method. In recent years, the discriminative method has also received extensive attention from researchers. The discriminative method is a method of distinguishing the object from the surrounding background based on considering various information of the object.

[0004] Over the years, through the experiments, explorations and researches of many scientists such as Bolme, Henriques, Danelljan, Ma, Galoogahi and Muller, a series of significant progress has been made in visual object tracking methods, but there are still the following problems:

[0005] 1. When performing visual object tracking, the object and surrounding background information are not considered. When the object is occluded or the environmental illumination changes, model drift is likely to occur, affecting the tracking performance of the object. At the same time, the boundaries of image sample blocks are discontinuous, and spatial boundary effects are likely to occur.

[0006] 2. The diversity and redundancy problems between different features of the object are not considered, which will affect the accuracy of object tracking.

[0007] 3. In the process of object tracking, traditional methods mostly directly use the maximum value of the object response to locate the object, and this kind of positioning method has low accuracy.

[0008] 4. In the process of object tracking, accidental peaks are likely to occur, resulting in model drift or tracking failure.

[0009] In summary, how to design an efficient and robust visual object tracking algorithm is still an urgent problem to be solved at present. Summary of the Invention

[0010] To solve the above problems, the present invention provides a robust visual object tracking algorithm from coarse to fine based on sparse learning, which can effectively suppress the influence of internal interferences such as target scale changes and external interferences such as background clutter and illumination changes, effectively avoid the occurrence of model drift or tracking failure caused by accidental peaks, and improve the positioning accuracy at the same time.

[0011] A robust visual object tracking algorithm from coarse to fine based on sparse learning includes the following steps:

[0012] S1: Establish a robust target appearance model;

[0013] S11: Introduce a regularization constraint on the basis of classical filtering:

[0014] S12: Select features with positive reference value from the target model using lasso regression, and introduce λ1||w||1;

[0015] S13: Add a regularization constraint: Embed the sparsity of the target response into the tracking framework;

[0016] S14: Obtain the model objective function in a single channel:

[0017]

[0018] where, x0 ∈ R n represents the extracted target, x i ∈ R n i ∈ [1,4] represents the surrounding background features, y ∈ R n represents the ideal two-dimensional Gaussian response, w ∈ R n represents the filter to be learned, the symbol * represents the circular convolution operation, the parameter λ1 is used to control the intensity of feature selection, the parameter λ2 is used to control the intensity of the background response regression to 0, and λ SR is used to control the sparsity degree of the response;

[0019] S15: Gradually optimize the objective function and solve the filter;

[0020] S2: Track the target;

[0021] S21: Input the video sequence and select the (t - 1)-th frame as the initial frame of the target;

[0022] S22: Determine the initial state of the initial frame, initialize the relevant parameters of the interactive Kalman, and solve the filter w of the initial frame through the parameters and the method in step S15 t-1 ;

[0023] S23: Select a search area at the center position according to the initial state of the target, and extract the target features;

[0024] S24: Based on the extracted target features, train the filter w t-1 to obtain the coarse positioning filter w t-1a and the fine positioning filter w t-1b ;

[0025] S25: According to w t-1a and w t-1b trained from the (t - 1)-th frame image, determine the relevant state of the target in the t-th frame by a coarse-to-fine positioning method;

[0026] S26: According to the determined target state in the t-th frame, train the filter w t for the t-th frame.

[0027] S27: Update the filter w t , and at the same time update the relevant parameters of the interactive Kalman filter;

[0028] S28: Iteratively solve to obtain w t+1 , w t+2 , w t+3 ......, until the optimal solution of w is obtained.

[0029] Preferably, when positioning the target by a coarse-to-fine positioning method in step S25, use the interactive Kalman filter to supervise the positioning of the target; when the target state positioned by the coarse-to-fine positioning method is not good, the interactive Kalman filter is used to position the target.

[0030] Preferably, step S24 includes a coarse positioning filter w t-1a , and step S25 includes the following coarse positioning steps:

[0031] S251: According to the target state determined in the (t - 1)-th frame, select the search area for coarse positioning;

[0032] S252: Extract the depth features of the target and perform a circular convolution operation with the coarse positioning filter w t-1 to obtain the response map for target coarse positioning;

[0033] S253: Determine the center position of the target according to the maximum value of the response map, and at the same time use the interactive Kalman filter to estimate the center position of the target;

[0034] S254: Perform a screening of the coarse positioning positions according to the two center positions determined in step S253, and at the same time use the target position determined by the interactive Kalman filter for supervision;

[0035] S255: After the rough positioning position is screened, if the target displacement determined by the depth feature and the rough positioning filter w t-1 is greater than several times the target displacement determined by the interactive Kalman filter, the rough positioning is performed using the center position estimated by the interactive Kalman filter; otherwise, the rough positioning is performed using the center position determined by the response map obtained by the depth feature and the rough positioning filter w t-1 .

[0036] Preferably, three fine positioning filters w are included in step S24 t-1b , and step S25 includes the following fine positioning steps:

[0037] S256: According to the rough positioning target position determined in step S255, determine the search window and extract the target features;

[0038] S257: Perform a circular convolution operation according to the extracted depth feature of the target and the fine positioning filter w t-1b to obtain the response map during target fine positioning;

[0039] S258: Perform adaptive fusion on the three response maps for fine positioning to obtain the final response value of fine positioning;

[0040] S259: Perform multi-peak detection on the response value;

[0041] S2510: If there are interferences of multiple unexpected peaks in the response map, the filter stops updating, and at the same time, the position of the target is estimated using the interactive Kalman; if the response map shows obvious peaks, the target is normally tracked; the target state of the t-th frame is obtained.

[0042] Preferably, S15 includes the following steps:

[0043] S151: According to the circular convolution theorem, obtain the objective function (2) based on the objective function (1):

[0044]

[0045] where the matrices X0, X i are the representations of all the movements of the target feature x0 and the background feature x i respectively;

[0046] S152: Optimize the objective function (2), introduce the equality constraint w = g, and obtain the objective function (3):

[0047]

[0048] S153: Incorporate the equality constraint w = g into the objective function (2) to obtain the augmented Lagrangian form of the objective function:

[0049]

[0050] Among them, s represents the Lagrange multiplier, and ρ > 0 represents the penalty coefficient of the quadratic term;

[0051] S154: Use the Alternating Direction Method of Multipliers (ADMM) to solve for w * 、g, and s through iteration:

[0052]

[0053] Preferably, given g and s, optimize and solve the sub-problem w * in step S154: Redefine the sub-formula

[0054]

[0055] in formula (5) as:

[0056]

[0057] Among them, A is the data matrix stacked by the target image block X0, its sparse regularization term, and the context image block X i and the specific forms of A and are as follows:

[0058]

[0059] Take the derivative of function (7), set the derivative equal to 0, and obtain the closed-form solution of the filter w:

[0060]

[0061] The data matrix A is a circulant matrix. Utilize the property that a circulant matrix can be diagonalized in the frequency domain to obtain the optimal solution of the filter w in the frequency domain:

[0062]

[0063] Preferably, given w and s, optimize and solve the sub-problem g in step S154. The optimization variable g is equivalent to optimizing the sub-formula

[0064]

[0065] Among them, Both λ1||g||1 are convex functions, is the locally differentiable part, and λ1||g||1 is the locally non-differentiable part. Through the shrinkage threshold operator, obtain the closed-form solution of g:

[0066]

[0067] Among them, max(a, b) represents taking the maximum value of a and b, and sign(a) is the sign function, and its specific form is as follows:

[0068]

[0069] Preferably, given w and g, optimize and solve the sub-problem s in step S154: solve the sub-formula of s in formula (5) through formula (13), and obtain s:

[0070]

[0071] Among them, ρ max represents the maximum value of the penalty parameter, and β represents the iteration step size.

[0072] Preferably, obtaining the final response value for fine positioning in step S258 includes the following steps:

[0073] S2581: According to empirical assumptions, the extracted depth features reduce the resolution of the image while representing the high-level information of the target. Conversely, the shallow features represent the low-level information of the image and increase the resolution of the image;

[0074] S2582: Assign different weights, i.e., [α1, α2, α3], to the three filtering templates during fine positioning and the response maps generated from the corresponding features of the three filtering templates, where α1 < α2 < α3;

[0075] S2583: Calculate the reliabilities β1, β2, β3 of the three fine positioning response maps. The reliability is characterized by the average peak energy, i.e.:

[0076]

[0077] S2584: The final weight of each layer of the response map is

[0078] S2585: Obtain the final response value for fine positioning as:

[0079]

[0080] Preferably, performing multi-peak detection on the response value in step S259 includes the following steps:

[0081] S2591: Perform a shift transformation on the response map R to move the zero-frequency component of the response map to the data center and rearrange the response map R, formulated as:

[0082] represents the circular shift operation;

[0083] S2592: Determine the peak and sub-peak of the response graph second = τ × peak, where τ ∈ [0, 1] is a constant;

[0084] S2593: Screen the peak and sub-peak second of the response values to form a list, which stores the response value r and the corresponding position (x r , y r ) in the response graph.

[0085] S2594: Set the response values in the response graph R that are less than the sub-peak second to 0 to obtain a new response graph R new ;

[0086] S2595: Traverse all the values in the list of the response graph R new to determine the number of interfering peaks; if the number of interfering peaks is 0, use the response graph R to normally perform target tracking, otherwise, use interactive Kalman filtering to track the target.

[0087] Advantages of the present invention:

[0088] 1. When establishing a robust target appearance model, the present invention integrates the target and its surrounding background information into the filter at the same time, which can suppress the influence of internal interferences such as boundary effects and external interferences such as occlusion and background clutter on the tracking performance.

[0089] 2. The present invention uses lasso regression to select features with positive reference value from the target model, which can avoid the adverse effects of the correlation and redundancy between target features and suppress the potential interference information in the original feature space.

[0090] 3. The present invention embeds the sparsity of the target response into the tracking framework, which can suppress the situation that when the target appearance model changes greatly frequently, the model drifts or even the tracking fails due to multiple peaks in the target response value.

[0091] 4. The present invention adopts a coarse-to-fine positioning method and simultaneously introduces interactive Kalman filtering to supervise and participate in the target positioning:

[0092] 1) The coarse-to-fine positioning method can improve the accuracy of target positioning;

[0093] 2) Interactive Kalman filtering can further supervise the accuracy of coarse positioning and fine positioning. When the states of coarse positioning and fine positioning are not good, interactive Kalman filtering is used to position the target; using Kalman filtering to position the target can further alleviate the drift of the model;

[0094] 3) Detecting multiple peaks in the response map during fine positioning can improve the accuracy of fine positioning. Description of the Drawings

[0095] Figure 1 is the flow chart of the target tracking algorithm in the present invention.

[0096] Figure 2 is the flow chart of the target tracking steps in the present invention.

[0097] Figure 3 is the schematic diagram of multiple peak detection in the present invention.

[0098] Figure 4 is the flow chart of the interactive Kalman filter in the present invention. Detailed Embodiments

[0099] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with the attached Figures 1-4 drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.

[0100] A robust visual target tracking algorithm from coarse to fine based on sparse learning includes the following steps:

[0101] S1: Establish a robust target appearance model;

[0102] S11: Integrate the internal interference and external interference information into the filter, and the formula is to introduce a regularization constraint on the basis of classical filtering: where the internal interference includes the target and its surrounding background information, boundary effects, etc., and the external interference includes occlusion, background clutter, etc. By introducing a regularization constraint on the basis of classical filtering, the influence of internal interference and external interference on the target tracking performance can be avoided;

[0103] S12: Select the features with positive reference value from the target model using lasso regression, and introduce λ1||w||1;

[0104] S13: Based on the correlation and redundancy between target features, in order to avoid the influence of potential interference information in the original feature space on the tracking performance, add a regularization constraint: Embed the sparsity of the target response into the tracking framework;

[0105] S14: Obtain the model objective function under a single channel:

[0106]

[0107] where, x0∈R n represents the extracted target, xi ∈R n i ∈ [1, 4] represents the surrounding background features, y ∈ R n represents the ideal two-dimensional Gaussian response, w ∈ R n represents the filter to be learned. The symbol * represents the circular convolution operation. The parameter λ1 is used to control the strength of feature selection, and the parameter λ2 is used to control the strength of the background response regressing to 0. λ SR is used to control the sparsity of the response;

[0108] S15: Gradually optimize the objective function and solve for the filter, including the following steps:

[0109] S151: Based on the circular convolution theorem, obtain the objective function (2) from the objective function (1):

[0110]

[0111] where the matrices X0, X i are the target feature x0 and the background feature x i respectively for all the shift representations;

[0112] S152: Optimize the objective function (2), introduce the equality constraint w = g, and obtain the objective function (3):

[0113]

[0114] S153: Incorporate the equality constraint w = g into the objective function (2) to obtain the augmented Lagrangian form of the objective function:

[0115]

[0116] where s represents the Lagrange multiplier, and ρ > 0 represents the penalty coefficient of the quadratic term, which is used to adjust the convergence degree;

[0117] S154: Further optimize the objective function (4), and use the alternating direction method of multipliers (ADMM) to solve for w * , g, and s by iteration to optimize the augmented Lagrange function:

[0118]

[0119] Optimize and solve the sub-problem w in step S154 * : Given g and s, combine the first three sub-formulas in formula (5)

[0120] and re-define it as:

[0121]

[0122] Among them, A is the data matrix stacked by the target image block X0, its sparse regularization term, and the context image block X i stacked, A and The specific form of is as follows:

[0123]

[0124] Derive the function (7), set the derivative equal to 0, and obtain the closed-form solution of the filter w:

[0125]

[0126] The data matrix A is a circulant matrix. Utilize the property that a circulant matrix can be diagonalized in the frequency domain to obtain the optimal solution of the filter w in the frequency domain:

[0127]

[0128] Optimize and solve the sub-problem g in step S154: Given w and s, the optimization variable g is equivalent to optimizing the sub-formula in formula (5)

[0129]

[0130] Among them, Both λ1||g||1 are convex functions, is the locally differentiable part, and λ1||g||1 is the locally non-differentiable part. Through the shrinkage threshold operator, the closed-form solution of g is obtained:

[0131]

[0132] Among them, max(a, b) represents taking the maximum value of a and b, and sign(a) is the sign function. The specific form is as follows:

[0133]

[0134] Optimize and solve the sub-problem s in step S154: Given w and g, solve the sub-formula about s in formula (5) through formula (13), and obtain s:

[0135]

[0136] Among them, ρ max represents the maximum value of the penalty parameter, and β represents the iteration step size.

[0137] S2: Track the target;

[0138] S21: Input the video sequence, and select the (t - 1)-th frame as the initial frame of the target;

[0139] S22: Determine the initial state such as the center position and range of the initial frame, initialize the relevant parameters of the interactive Kalman, and solve for the filter w of the initial frame through the parameters and the above method t -1;

[0140] S23: Select a search area at the center position according to the initial state of the target. The search area is 2 times the target scale; extract the target features, including the convolutional features (pool1 layer, pool3 layer, and Relu5-2) extracted by the VGG19 network and the manual features (HOG, CN, and gray);

[0141] S24: Based on the extracted target features, train the filter w t-1 Obtain four filters, including a coarse localization filter w t-1a (Relu5-2 layer) and three fine localization filters w t-1b ;

[0142] S25: According to w t-1a and w t-1b trained from the (t - 1)-th frame image, determine the target center position, scale, and the coefficients of the Kalman filter through a coarse-to-fine localization method, and determine the relevant state of the target in the t-th frame; when localizing the target through a coarse-to-fine localization method, use the interactive Kalman filter to supervise the localization of the target; when the target state localized by the coarse-to-fine localization method is not good, use the interactive Kalman filter to localize the target.

[0143] Coarse localization steps:

[0144] S251: According to the target state determined from the (t - 1)-th frame, select the search area for coarse localization. Here, the search area is 2.5 times the target scale;

[0145] S252: Perform a circular convolution operation on the target depth features (Relu5-2 layer) extracted by the VGG19 network and the coarse localization filter w t-1 to obtain the response map for target coarse localization;

[0146] S253: Determine the center position of the target according to the maximum value of the response map, and at the same time use the interactive Kalman filter to estimate the center position of the target;

[0147] S254: Perform a screening of the coarse localization positions according to the two center positions determined in step S253, and at the same time use the target position determined by the interactive Kalman filter for supervision;

[0148] S255: After the screening of the coarse localization positions, if after the depth features and the coarse localization filter w t-1If the determined target displacement is greater than several times the target displacement determined by the interactive Kalman filter, rough positioning is performed using the central position estimated by the interactive Kalman filter; otherwise, the depth feature and the rough positioning filter w t-1 Rough positioning is performed using the central position determined from the obtained response map.

[0149] Fine positioning steps:

[0150] S256: According to the rough positioning target position determined in step S255, a search window is determined and target features are extracted at 7 scales, where the search window is 2 times the target scale, and the target features include CNN features and handcrafted features;

[0151] S257: According to the depth feature of the extracted target and the fine positioning filter w t-1b Perform a circular convolution operation to obtain the response map during target fine positioning;

[0152] S258: Adaptive fusion is performed on the three response maps for fine positioning to obtain the final response value for fine positioning; the depth feature can represent the semantic information of the image. Although the pooling operation increases the receptive field, it weakens the target localization ability. On the contrary, the shallow features and their handcrafted features can characterize the texture and contour information of the target. Therefore, after obtaining the response map during fine positioning, we need to perform adaptive fusion on the response map.

[0153] The adaptive fusion includes the following steps:

[0154] S2581: According to empirical assumptions, the depth features extracted by the CNN reduce the image resolution while representing the high-level information of the target. On the contrary, the shallow features represent the low-level information of the image and increase the image resolution;

[0155] S2582: Different weights, i.e., [α1, α2, α3], are assigned to the three filtering templates for fine positioning and the response maps generated from the corresponding features of the three filtering templates, where α1 < α2 < α3;

[0156] S2583: Calculate the reliability β1, β2, β3 of the three fine positioning response maps. The reliability is characterized by the average peak energy, i.e.:

[0157]

[0158] S2584: The final weight formula for each layer of the response map is Substitute [α1, α2, α3] in step S2582 and β1, β2, β3 in step S2583 into the formula to obtain the final weight;

[0159] S2585: The final response value for fine positioning is obtained as:

[0160]

[0161] S259: Perform multi-peak detection on the response value to determine whether there are multiple peaks in the response value;

[0162] The multi-peak detection includes the following steps:

[0163] S2591: Perform a shift transformation on the response map R to move the zero-frequency component of the response map to the data center and rearrange the response map R, formulated as:

[0164] Indicates a cyclic shift operation;

[0165] S2592: Determine the peak peak and sub-peak peak of the response map second = τ × peak, where τ ∈ [0, 1] is a constant;

[0166] S2593: Screen the response values between the peak peak and the sub-peak peak to form a list, and store the response value r and its corresponding position (x second in the response map in the list r , y r ).

[0167] S2594: Set the response values in the response map R that are less than the sub-peak peak second to 0 to obtain a new response map R new ;

[0168] S2595: The response map R new traverses all the values in the list to determine the number of interference peaks; if the number of interference peaks is 0, the target is tracked normally using the response map R, otherwise, the interactive Kalman filter is used to track the target;

[0169] S2510: If there are interferences from multiple unexpected peaks in the response map, the filter stops updating, and at the same time, the interactive Kalman filter is used to estimate the position of the target; if the response map shows obvious peaks, the target is tracked normally; obtain the target state of the t-th frame.

[0170] S26: According to the determined target state of the t-th frame, train the filter w t ;

[0171] S27: Update the filter w t in a moving average manner, that is: w t = αw + (1 - α)w t-1 , and at the same time update the relevant parameters of the interactive Kalman;

[0172] S28: Iteratively solve according to the above steps to obtain w t+1 and w t+2 and w t+3 ........, until the optimal solution of w is obtained.

[0173] It is worth mentioning that: Kalman filtering is established in the time domain space and can estimate the state of the system at the next moment based on the state of the system at the previous moment and the observation value at the current moment; considering that in actual target tracking, the target may experience various complex motions, such as: uniform linear motion, uniformly accelerated linear motion, and turning motion, it is inaccurate to use only a single dynamic model to characterize the motion state of the target. At this time, an integration of multiple dynamic models is required to characterize the motion state of the target. The interactive multiple model Kalman (IMM) can cover all possible motion states of the target; the basic idea of IMM is to use multiple filters to estimate different motion states of the target, then, each model implements prediction and correction in parallel, and finally, the corrected state estimate is obtained by weighting; among them, it is assumed that the probability transition between each model follows a Markov process.

[0174] Assume that the target has N different motion states, and the state equation and observation equation of the i-th model are as follows:

[0175] X i (t + 1) = A i (t)X i (t) + W i (t)

[0176] Z i (t + 1) = H i (t + 1)X i (t + 1) + V i (t + 1)

[0177] Among them, X i (t + 1) and Z i (t + 1) respectively represent the state and observation vector of the target at t + 1, A i (t) represents the state transition matrix of the i-th model, H i (t + 1) represents the observation matrix, W i (t) and V i (t + 1) respectively represent Gaussian white noise sequences. The probability transition between each model follows a Markov process, where the element p ij in the matrix P represents the probability of the i-th model transitioning to the j-th model, and the probability transition matrix is as follows:

[0178]

[0179] such as Figure 3As shown, after the system initialization, the IMM algorithm mainly includes the following four parts: (a) input interaction, (b) Kalman filtering, (c) model probability update, (d) output interaction;

[0180] The initial state of the interactive Kalman is determined by the first several frames of the video sequence;

[0181] Estimate the state of each model According to the state of each model and the model probability u i (k - 1) to obtain the mixed estimate after interaction and the covariance matrix Kalman filtering, based on the mixed state determined in the previous step and the covariance matrix as well as the observation value Z(t) as the current initial state to perform Kalman filtering to update the predicted state and the covariance matrix

[0182] Update the model probability u i (k);

[0183] Obtain the state of the target at the current moment and the covariance matrix

[0184] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0185] The above specific embodiments of the present invention do not constitute a limitation to the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included within the protection scope of the claims of the present invention.

Claims

1. A robust visual object tracking algorithm from coarse to fine based on sparse learning, characterized in that, It includes the following steps: S1: Establish a robust target appearance model; S11: Introduce a regularization constraint on the basis of classical filtering: S12: Select features with positive reference value from the target model using lasso regression, and introduce λ1||w||1; S13: Add a regular constraint: Embed the sparsity of the target response into the tracking framework; S14: Obtain the model objective function in a single channel: where \(x_0\in\mathbb{R}\) n denotes the extracted target, \(x\) i \(\in\mathbb{R}\) n for \(i\in[1,4]\) denotes the surrounding background features, \(y\in\mathbb{R}\) n denotes the ideal two-dimensional Gaussian response, \(w\in\mathbb{R}\) n denotes the filter to be learned, the symbol \(*\) represents the circular convolution operation, the parameter \(\lambda_1\) is used to control the strength of feature selection, the parameter \(\lambda_2\) is used to control the strength of the background response regressing to 0, \(\lambda\) SR is used to control the sparsity of the response; S15: Gradually optimize the objective function and solve for the filter; S2: Track the target; S21: Input the video sequence and select the (t - 1)-th frame as the initial frame of the target; S22: Determine the initial state of the target, initialize the relevant parameters of the interactive Kalman, and solve for the filter w of the initial frame through the parameters and the method in step S15 t-1 ; S23: Select a search area at the center position according to the initial state of the target and extract the target features; S24: Train the filter w based on the extracted target features t-1 to obtain the coarse localization filter w t-1a and the fine localization filter w t-1b ; S25: w trained based on the (t - 1)-th frame image t-1a and w t-1b , determine the relevant state of the target in the t-th frame through a coarse-to-fine positioning method; S26: Train the filter w for the t-th frame based on the determined target state of the t-th frame t ; S27: Update the filter w t , and update the relevant parameters of the interactive Kalman at the same time; S28: Iteratively solve to obtain w t+1 and w t+2 and w t+3 ...... until the optimal solution of w is obtained.

2. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 1, wherein In step S25, when positioning the target by a coarse-to-fine positioning method, interactive Kalman filtering is used to supervise the positioning of the target; when the target state positioned by the coarse-to-fine positioning method is not good, interactive Kalman filtering is used to position the target.

3. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 2, characterized in that: The step S24 includes a coarse positioning filter w t-1a , and the step S25 includes the following coarse positioning steps: S251: Select the search area for coarse positioning according to the target state determined in the (t - 1)-th frame; S252: Extract the depth features of the target and the coarse localization filter w t-1a Perform a circular convolution operation to obtain the response map during the coarse localization of the target; S253: Determine the center position of the target according to the maximum value of the response map, and at the same time use interactive Kalman filtering to estimate the center position of the target; S254: Screen the coarse positioning positions according to the two center positions determined in step S253, and at the same time use the target position determined by interactive Kalman filtering for supervision; S255: After the rough positioning position is screened, if the target displacement determined by the depth feature and the rough positioning filter w t-1a is greater than several times the target displacement determined by the interactive Kalman filter, the rough positioning is performed using the center position estimated by the interactive Kalman filter; otherwise, the rough positioning is performed using the center position determined by the response map obtained by the depth feature and the rough positioning filter w t-1a obtained.

4. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 3, characterized in that: Three fine positioning filters w are included in the step S24 t-1b , and the step S25 includes the following fine positioning steps: S256: Determine the search window and extract the target features according to the coarse positioning target position determined in step S255; S257: Perform a circular convolution operation based on the depth features of the extraction target and the fine localization filter w t-1b to obtain the response map during the fine localization of the target; S258: Adaptively fuse the three response maps for fine positioning to obtain the final response value of fine positioning; S259: Perform multi-peak detection on the response value; S2510: If there are interferences from multiple unexpected peaks in the response map, the filter stops updating, and at the same time use interactive Kalman to estimate the position of the target; if the response map shows obvious peaks, perform normal tracking on the target; Obtain the target state of the t-th frame.

5. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 4, wherein: The step S15 includes the following steps: S151: According to the cyclic convolution theorem, obtain the objective function (2) based on the objective function (1) Among them, the matrices X0 and X i are the target feature x0 and the background feature x i respectively; all the movement representations S152: Optimize the objective function (2), introduce the equality constraint w = g, and obtain the objective function (3): S153: Incorporate the equality constraint w = g into the objective function (2) to obtain the augmented Lagrangian form of the objective function: where s represents the Lagrange multiplier, and ρ > 0 represents the penalty coefficient of the quadratic term; S154: Solve for w, g, and s iteratively using the alternating direction method of multipliers (ADMM). * , and s:

6. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 5, characterized in that: Given g and s, optimize and solve the sub-problem w in step S154 * : Optimize and solve the sub-formula in formula (5) Redefine as: Among them, A is a data matrix formed by stacking the target image block X0, its sparse regularization term, and the context image block X i stacked together, and the specific forms of A and are as follows: Take the derivative of the function (7), set the derivative equal to 0, and obtain the closed-form solution of the filter w: The data matrix A is a circulant matrix. Utilize the property that a circulant matrix can be diagonalized in the frequency domain to obtain the optimal solution of the filter w in the frequency domain:

7. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 6, characterized in that: Given w and s, optimize and solve the sub-problem g in step S154. Optimizing the variable g is equivalent to optimizing the sub-formula in formula (5) Among them, both λ1||g||1 are convex functions, is the locally differentiable part, and λ1||g||1 is the locally non-differentiable part. Through the shrinkage threshold operator, a closed-form solution of g is obtained: where max(a, b) represents taking the maximum value of a and b, and sign(a) is the sign function, and the specific form is as follows:

8. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 7, characterized in that: Given w and g, optimize and solve the sub-problem s in step S154: Solve the sub-formula about s in formula (5) through formula (13) and obtain s: where ρ max represents the maximum value of the penalty parameter, and β represents the iteration step size.

9. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 8, characterized in that: The step of obtaining the final response value of fine positioning in step S258 includes the following steps: S2581: According to the empirical hypothesis, the extracted deep features reduce the resolution of the image while representing the high-level information of the target. Conversely, the shallow features represent the low-level information of the image and increase the resolution of the image; S2582: Different weights, i.e., [α1, α2, α3], are assigned to the three filtering templates during fine localization and the response maps generated from the corresponding features of the three filtering templates, where α1 < α2 < α3; S2583: Calculate the reliabilities β1, β2, and β3 of the three fine localization response maps. The reliability is characterized by the average peak energy, i.e.: S2584: The final weight of each layer's response map is S2585: The final response value for fine localization is obtained as:

10. The robust visual object tracking algorithm from coarse to fine based on sparse learning according to claim 9, characterized in that: The multi-peak detection of the response value described in step S259 includes the following steps: S2591: Perform a shift transformation on the response map R to move the zero-frequency component of the response map to the data center and rearrange the response map R, formulated as: Indicates a cyclic shift operation; S2592: Determine the peak and sub-peak of the response graph second = τ × peak, where τ ∈ [0, 1] is a constant; S2593: Screen the peak value peak and the secondary peak value peak second The response values between are formed into a list, and the response value r and the corresponding position (x r , y r ) in the response graph are stored in the list; S2594: Set the response values in the response graph R that are less than the secondary peak peak second to 0 to obtain a new response graph R new ; S2595: Response Diagram R new Traverse all the values in the list to determine the number of interference peaks; if the number of interference peaks is 0, perform target tracking normally using the response diagram R, otherwise, use interactive Kalman filtering to track the target.