A Correlated Particle Filter Video Tracking Method with Multi-Feature Cascade Fusion

Through the related particle filtered video tracking method of multi-feature cascade fusion, the robustness problem of video target tracking in complex backgrounds is solved, and high-precision and efficient video target tracking is achieved.

CN114372998BActive Publication Date: 2025-07-08CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111529886.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-07-08
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The existing video target tracking algorithm is not robust enough in complex backgrounds such as lighting changes, target occlusion and scale changes. Single feature detection is susceptible to interference, resulting in inaccurate tracking.

Method used

The video tracking method of related particle filtering with multi-feature cascade fusion is adopted. Through initialization settings, particle motion prediction, multi-level feature update and fusion processing, combined with the target's dynamic state, color and edge characteristics, the correlation filter is used to improve tracking accuracy and robustness.

Benefits of technology

Maintain high-precision tracking in complex contexts, reduce computing complexity, improve real-time and robustness, and enhance the reliability of video target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003410405560000061
    Figure BDA0003410405560000061
  • Figure BDA0003410405560000062
    Figure BDA0003410405560000062
  • Figure BDA0003410405560000069
    Figure BDA0003410405560000069
Patent Text Reader

Abstract

The present invention discloses a correlated particle filtering video tracking method with multi-feature cascade fusion, aiming at a series of challenging problems such as target deformation and occlusion in the visual tracking process and the problem of weak robustness of single-feature target tracking. Using the algorithm framework of correlated particle filtering, the correlation filter guides the sampled particles to the mode of the target state distribution. At the same time, a series connection hierarchical fusion method and a two-stage particle weight update method are adopted. Thereby reducing the computational complexity of particle filtering and improving the real-time performance and accuracy of visual tracking. It can still track the target with high accuracy under the influence of challenging factors such as occlusion, illumination change, and scale transformation, showing stronger robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video object tracking method, and more particularly to a related particle filter video tracking method with multi-feature cascade fusion. Background Art

[0002] Video object tracking is one of the core research topics in the field of computer vision and has extremely wide applications in many aspects such as video surveillance, human-computer interaction, visual navigation, and medical diagnosis. Video object tracking technology refers to detecting and identifying an object of interest in subsequent video sequence images based on a previously selected object of interest, and obtaining motion parameters such as the position and size of the object. In actual application scenarios, there are complex background situations such as illumination changes, object occlusion, and object deformation, resulting in missing object information or interference in obtaining object information, thus leading to inaccurate tracking. In order to ensure the reliability of the object tracking effect, it is necessary to further improve the robustness of the object tracking algorithm.

[0003] Particle filter realizes recursive Bayesian filtering through a non-parametric Monte Carlo method. The idea of particle filter is to use a large number of particles to simulate the state distribution of the object. The larger the number of particles, the more accurate the final position estimation of the object. However, a large number of particles increases the computational complexity of particle filter. The object tracking algorithm based on correlation filtering introduces the idea of correlation filtering into object tracking applications. By designing a correlation filter template, when it acts on the tracking object, the obtained response is the largest, and the position of the maximum response value is used as the estimated position of the object, making the object tracking more robust. However, the tracker based on correlation filtering still cannot handle complex backgrounds such as scale changes and object occlusion well. In recent years, the related particle filter algorithm has been proposed, which combines correlation filtering and particle filtering to improve the accuracy and robustness of the filter when tracking the object. Using only the detection of a single feature as the observation value of the object is often unreliable and is easily affected by external or self-interference. Aiming at a series of challenging problems such as object deformation and occlusion in the video object tracking process and the weak robustness of single-feature object tracking, the present invention proposes a video object tracking algorithm with multi-feature cascade fusion under the framework of the related particle filter algorithm. Summary of the Invention

[0004] Aiming at a series of challenging problems such as object deformation and occlusion in the visual tracking process and the weak robustness of single-feature object tracking, the present invention proposes a visual tracking algorithm with multi-feature cascade fusion under the framework of the related particle filter algorithm.

[0005] To achieve the above tasks, the present invention adopts the following technical solutions:

[0006] A related particle filtering video tracking method based on multi-feature cascade fusion, the method comprising the following steps:

[0007] Step 1: Obtain a set of video images containing M frames, and perform initialization settings on the initial frame image, including initializing the dynamic state of the centroid of the target corresponding to the initial frame image, initializing the centroid particle set, and initializing the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image;

[0008] Step 2: Read the frame image corresponding to the next frame, and predict the particle motion state of the current frame image according to the centroid dynamic motion state of the particles corresponding to the previous frame image, so as to obtain the predicted state of the particle set of the current frame image;

[0009] Step 3: Perform a first-level update on the predicted state of the particle set of the current frame image according to the target color feature correlation filter corresponding to the previous frame image;

[0010] Step 4: Perform a second-level update on the particle set state of the current frame image obtained in Step 3 according to the target edge feature correlation filter corresponding to the previous frame image;

[0011] Step 5: Process the particle set state of the current frame image obtained in Step 4 to obtain the centroid dynamic estimation state of the current frame target;

[0012] Step 6: Based on the centroid dynamic estimation state of the current frame target obtained in Step 5, update the target color feature, the target edge feature, and the corresponding target color feature correlation filter and target edge feature correlation filter of the current frame image;

[0013] Step 7: Repeat Step 2 to Step 6 until each frame image in the M-frame video image has completed the above processing and then end.

[0014] Furthermore, the present invention further includes the following technical features:

[0015] In Step 1, the dynamic state of the centroid of the target corresponding to the initial frame image is initialized, and the dynamic equation of the centroid of the target is:

[0016] x k = Ax k-1 + ω k

[0017] where A is the state transition matrix of the centroid motion of the target, T is the sampling time interval, and ω k is Gaussian white noise with a mean of zero; the initial dynamic state of the centroid of the target is set as x0 and y0 respectively represent the horizontal and vertical coordinates of the centroid particle of the target, They respectively represent the moving speeds of the target centroid particles in the x-direction and y-direction.

[0018] Further, the initialization setting process of step 1 for the target centroid particle set is as follows: Set a target bounding box centered on the target centroid, use the initial dynamic state of the target centroid as the mean of the normal distribution, construct the covariance of the normal distribution with the diagonal of the target bounding box, randomly sample the target bounding box to obtain N random particles, and assign equal weights to the N particles.

[0019] Further, the initialization setting process of step 1 for the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image is as follows: Centered on the centroid of the target bounding box, select a window with a size 2.5 times that of the target bounding box, use the cyclic shift method to construct a training sample set for this window, extract the color features and edge features of the training sample set, and use the ridge regression method to train the color features and edge features respectively to obtain the target color feature correlation filter and the target edge feature correlation filter.

[0020] Step 3 performs a first-level update on the predicted state of the particle set of the current frame image according to the target color feature correlation filter corresponding to the previous frame image. The specific process is as follows:

[0021] Step 3.1: Respectively centered on the position of each particle in the current frame image, select a search region with a size 2.5 times that of the target bounding box, use the cyclic shift method to construct a sample set for this search region, extract the color features of the sample set, and obtain the color features of each particle in the current frame image.

[0022] Step 3.2: Successively perform window function processing and fast Fourier transform on the color features of each particle in the current frame image obtained in step 3.1 to obtain the processed color features, multiply the processed color features by the color feature correlation filter corresponding to the previous frame image, and perform the inverse fast Fourier transform on the multiplication result to obtain the target color feature correlation response map of each particle in the current frame image.

[0023] Step 3.3: Use the positions of the maximum response values of the target color feature correlation response maps of each particle in the current frame image obtained in step 3.2 as the new positions of the corresponding particles in the current frame image respectively, use the maximum response values of the target color feature correlation response maps of each particle as the weight values of the corresponding particles in the current frame image, and then perform normalization processing on the obtained particle weight values to obtain the estimated state of the particle set of the current frame image after the first-level update.

[0024] Step 4 performs a second-level update on the estimated state of the particle set of the current frame image obtained in step 3 according to the target edge feature correlation filter corresponding to the previous frame image. The specific process is as follows:

[0025] Step 4.1: Centered at the position of each particle in the current frame image obtained in Step 3, a search region with a size 2.5 times that of the target bounding box is selected. The sample set is constructed by using the circular shift method for this search region, and the edge features of the sample set are extracted to obtain the edge features of each particle in the current frame image.

[0026] Step 4.2: The edge features of each particle in the current frame image obtained in Step 4.1 are sequentially subjected to window function processing and fast Fourier transform to obtain the processed edge features. The processed edge features are multiplied by the edge feature correlation filter corresponding to the previous frame image, and the multiplication result is subjected to inverse fast Fourier transform to obtain the target edge feature correlation response map of each particle in the current frame image.

[0027] Step 4.3: The positions of the maximum response values of the target edge feature correlation response maps of each particle in the current frame image obtained in Step 4.2 are respectively used as the positions of the corresponding particles in the current frame image, and the maximum response value of the target edge feature correlation response map is used as the weight value of the corresponding particle in the current frame image. Then, the obtained particle weight values are normalized to obtain the estimated state of the particle set of the current frame image after the secondary update.

[0028] Further, Step 6 is premised on the target centroid dynamics estimation state of the current frame image obtained in Step 5, and updates the target color feature correlation filter and the target edge feature correlation filter corresponding to the current frame image, which is realized by the following formula:

[0029]

[0030]

[0031] Where: is the updated target color feature correlation filter of the current frame (the kth frame), is the target color feature correlation filter of the previous frame (the k - 1th frame) image, is the target color feature correlation filter of the current frame (the kth frame) before update; is the updated target edge feature correlation filter of the current frame (the kth frame), is the target edge feature correlation filter of the previous frame (the k - 1th frame) image, is the target edge feature correlation filter of the current frame (the kth frame) before update, and η is the learning factor η ∈ [0, 1].

[0032] Premised on the target centroid dynamics estimation state of the current frame image obtained in Step 5, the target color feature and the target edge feature corresponding to the current frame image are updated, which is realized by the following formula:

[0033]

[0034]

[0035] Among them, The target color feature corresponding to the updated current frame (the k-th frame) image The target color feature of the previous frame (the (k - 1)-th frame) image The target color feature corresponding to the current frame (the k-th frame) image before update The target edge feature of the updated current frame (the k-th frame) image The target edge feature corresponding to the previous frame (the (k - 1)-th frame) image The target edge feature corresponding to the current frame (the k-th frame) image before update, where η is the learning factor and η ∈ [0, 1].

[0036] The present invention has the following technical features compared with the prior art:

[0037] (1) Within the framework of the particle filter, the present invention uses a correlation filter to guide the particles to the region with a higher probability of the target's existence, thereby reducing the number of redundant particles, reducing the computational complexity of the particle filter, and improving the real-time performance of visual tracking.

[0038] (2) The present invention uses a multi-feature cascade fusion method to perform multi-level updates on the target position and fuses the multi-level estimated positions, thereby improving the accuracy of visual tracking.

[0039] (3) Under the influence of challenging factors such as occlusion, illumination change, and scale transformation, the present invention can still track the target with high accuracy, showing stronger robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flow chart of the overall method of the present invention.

[0041] Figure 2 is a comparison chart of the success rate algorithms of the present invention and the prior art. Among them, Figure 2a OPE success rate curve graph, Figure 2b SRE success rate curve graph, Figure 2c TRE success rate curve graph;

[0042] Figure 3 is a comparison chart of the tracking accuracy of the present invention and the prior art. Among them, Figure 3a OPE tracking accuracy curve graph, Figure 3b SRE tracking accuracy curve graph, Figure 3c TRE tracking accuracy curve graph;

[0043] Figure 4 is a comparison chart of the success rate of the present invention and the prior art under different attributes, where: Figure 4a Comparison chart of the success rate under the attribute of light change; Figure 4b Comparison chart of the success rate under the attribute of scale change;

[0044] Figure 4c Comparison chart of the success rate under the occlusion attribute. Specific implementation manners

[0045] The technical content of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0046] See Figure 1 , a related particle filter video tracking method for multi-feature cascade fusion of the present invention, the method includes the following steps:

[0047] Step 1: Obtain a set of video images containing M frames, and perform initialization settings on the initial frame image, including initialization settings for the dynamic state of the target centroid corresponding to the initial frame image, initialization settings for the dynamic state of the centroid particle set, and initialization settings for the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image.

[0048] Step 2: Read the frame image corresponding to the next frame, and predict the particle motion state of the frame image corresponding to the current frame according to the dynamic motion state of the particles corresponding to the previous frame image, so as to obtain the predicted state of the particle set of the frame image corresponding to the current frame.

[0049] Step 3: Perform a first-level update on the predicted state of the particle set of the current frame image according to the target color feature correlation filter corresponding to the previous frame image, so as to obtain the estimated state of the particle set of the current frame image.

[0050] Step 4: Perform a second-level update on the estimated state of the particle set of the current frame image obtained in Step 3 according to the target edge feature correlation filter corresponding to the previous frame image, so as to obtain the estimated state of the particle set of the current frame image after the second-level update.

[0051] Step 5: Process the estimated state of the particle set in the current frame image obtained in Step 4 to obtain the estimated state of the target centroid dynamics of the current frame image;

[0052] Step 6: Based on the estimated state of the target centroid dynamics of the current frame image obtained in Step 5, update to obtain the target color feature, the target edge feature, and the corresponding target color feature correlation filter and target edge feature correlation filter corresponding to the current frame image, as the input for processing the next frame image;

[0053] Step 7: Repeat Step 2 - Step 6 until the processing of each frame image in the M-frame video image is completed.

[0054] The following will explain the technical terms appearing in the above solution to help better understand the technical content of this application.

[0055] The target centroid dynamics motion model assumes that the dynamics state vector of the target centroid in the (k - 1)-th frame video image is x k-1 , y k-1 representing the coordinates of the target centroid, respectively representing the velocities of the target centroid in the x-axis direction and the y-axis direction.

[0056] The dynamics equation of the target centroid motion can be described as

[0057] x k = Ax k-1 + ω k (1)

[0058] where A is the state transition matrix. Assuming that the centroid of the target moves in a uniform straight line, then

[0059]

[0060] T is the sampling time interval, and ω k is Gaussian white noise with zero mean.

[0061] The target bounding box to be tracked refers to the target bounding box manually calibrated on the initial frame image.

[0062] The target color feature is represented by a color histogram. To extract the color histogram feature using the HSV space, it is first necessary to convert the RGB image into an HSV image, and then quantize the three components of H, S, and V respectively. According to people's perception of colors, the quantization method adopted here is to divide the H space with a value range of 0 - 360 into 8 unequal intervals, and divide the S and V spaces with a value range of 0 - 1 into 3 unequal intervals respectively, as shown in the following formula:

[0063]

[0064]

[0065]

[0066] After quantizing the three components of H, S, and V respectively, where H represents hue, measured in degrees, with a range of 0 - 360; S represents saturation, and V represents brightness, both with a range of 0 - 1. Then, a global color histogram representing the image is obtained by weighted summation of the three components. The process is as follows: Construct a one-dimensional feature vector from the quantized three components of H, S, and V: U = HQ s Qv +SQ v +V, where Q s and Q v are the quantization levels of the S component and the V component respectively. Take Q s = 3, Q v = 3, then U = 9H + 3S + V. According to the above formula, the number of bins U of the color histogram can be obtained: U = 7×9 + 2×3 + 2 = 71, and a one-dimensional histogram with 72 bins is obtained, that is, the value range of U is [0, 1,..., 71].

[0067] Suppose the target is an upright rectangular tracking frame with axes h x and h y . When extracting the color features of the target, to increase the reliability of the color distribution, we use the kernel function to assign values to the target pixels, so that the weight of the pixel value closer to the target center is greater. Then the HSV color histogram at the rectangular frame l of the target image is expressed as The calculation formula is:

[0068]

[0069] where is the length of the diagonal of the rectangular tracking frame, δ is the Kronecker delta function, M is the total number of pixels in the target image area, l m is the position of the m-th pixel in the target image area, l is the target center, b(l m ) represents the color quantization value of the pixel at l m in the target area, and C is the normalization constant:

[0070]

[0071] K is the Gaussian kernel function, defined as:

[0072] K(r) = exp{-r 2 / (2σ 2 )} (7)

[0073] where r represents the distance between a certain pixel point in the target area and the target center point, and the parameter σ is the radius of the kernel function (take σ = 1).

[0074] For target edge feature extraction, the Canny edge detection operator is used to obtain the edge gradient magnitude and direction, and a gradient histogram is established according to the direction angle of the edge pixel points, thereby obtaining the edge gradient histogram feature.

[0075] First, the Gaussian low-pass filter is used to smooth the target image. Then, the Canny edge detection operator is applied to the smoothed target image to calculate the gradient of each pixel, obtaining the magnitude G(x, y) and direction angle θ(x, y) of the edge gradient:

[0076]

[0077] Among them, G x (x, y), G y (x, y) represent the gradient magnitudes in the horizontal and vertical directions respectively, and x, y represent the coordinates of any pixel point in the image. Here, the gradient magnitudes in the horizontal and vertical directions are approximately calculated using a 2×2 first-order finite difference:

[0078]

[0079]

[0080] In the above formula, f s (x, y) represents the pixel value of any pixel point in the smoothed target image.

[0081] Finally, a gradient histogram is constructed using the gradient direction angle of each pixel in the image. Here, a 36-bit direction gradient histogram is adopted, which is evenly divided into 36 regions with 10 degrees as a unit, and an edge-weighted gradient direction histogram is constructed with the target center l

[0082]

[0083] Among them, is the length of the diagonal of the rectangular tracking box, δ is the Kronecker delta function, M is the total number of pixels in the target image region, l m is the position of the m-th pixel in the target image region, l is the target center, G(l m ) represents the gradient magnitude of the pixel at l m in the target region, b(l m ) represents the quantization index value corresponding to the gradient direction angle of the pixel at l m in the target region, and C is the normalization constant (see formula (6)).

[0084] The target edge feature extraction uses the Canny edge detection operator to obtain the magnitude and direction of the edge gradient, and a gradient histogram is established based on the direction angle of the edge pixel points, thereby obtaining the edge gradient histogram feature.

[0085] The algorithm framework adopted by the present invention is the correlated particle filter algorithm. The particle filter is used to predict the motion state of the target, and the correlation filter is used to guide the particles to the maximum probability local area of the target probability distribution, making full use of the prior information and posterior information of the target. The multi-feature information cascade fusion method is adopted to perform multi-level updates of the particle positions, and multi-level fusion is performed on the estimated target positions.

[0086] Specifically, step 1: The initialization settings for the initial frame image are as follows:

[0087] For the initial frame k = 0, manually calibrate the target bounding box to be tracked in the initial frame image, and record the initial dynamic state of the target centroid as x0 and y0 respectively represent the coordinates of the target centroid, respectively represent the motion speeds of the target in the x direction and y direction.

[0088] Initialize the target centroid particle set. Calibrate the target bounding box with the target centroid position as the center, use the center position of the target bounding box as the mean of the normal distribution, construct the covariance of the normal distribution with the diagonal of the target bounding box, and randomly sample N particles from the normal distribution N(x0, Q) to obtain the particle dynamic state where x0 represents the mean of the normal distribution, which is the initial dynamic state of the target centroid manually calibrated, and Q = diag(Q 11 , Q 22 , Q 33 , Q 44 ) is the selected diagonal covariance matrix, and the diagonal elements take appropriately large values. Assign equal weights to each particle to obtain the particle set of the initial dynamic state of the target centroid in the initial frame

[0089] Initialize the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image. Select a window with a size 2.5 times that of the target bounding box with the target centroid as the center, extract the color feature and edge feature respectively, construct the training samples using circular shift based on the kernel correlation filter, and use the ridge regression method to train a target color feature correlation filter and a target edge feature correlation filter

[0090] Furthermore, regarding step 3, perform a first-level update on the predicted state of the particle set corresponding to the current frame image according to the target color feature correlation filter corresponding to the previous frame image. The specific process is as follows:

[0091] Step 3.1: For each particle Centered at its position coordinates, a search region 2.5 times the size of the target bounding box is selected. Based on the kernel correlation filter, for each predicted search region, color histogram features are extracted. After these features pass through the cosine window function, a fast Fourier transform is performed, and then they are correlated with the target color feature correlation filter corresponding to the previous frame multiplied, and the result of the multiplication is subjected to an inverse fast Fourier transform to obtain the target color feature correlation response map.

[0092] Step 3.2: Let the position where the maximum value of the target color feature correlation response map is located be Translate the current particle to a new position and use the maximum response value of the response map as the weight of the particle and normalize the weights

[0093] Let the particle set updated by the first-level color feature correlation filter be

[0094] Furthermore, regarding Step 4: The state of the particle set of the current frame image obtained in Step 3 is updated at the second level according to the target edge feature correlation filter corresponding to the previous frame image. The specific process is as follows:

[0095] Step 4.1: For the particle set obtained in Step 3 Centered at the position coordinates of the particle state a search region 2.5 times the size of the target bounding box is selected again. Based on the kernel correlation filter, for each predicted search region, edge feature vectors are extracted. After the feature vectors pass through the cosine window function, a fast Fourier transform is performed, and then they are correlated with the target edge feature correlation filter corresponding to the previous frame multiplied, and the result of the multiplication is subjected to an inverse fast Fourier transform to obtain the target edge feature correlation response map.

[0096] Step 4.2: Let the position where the maximum value of the target edge feature correlation response map is located be Translate the current particle to a new position and use the maximum response value of the response map as the weight of the particle and normalize the weights

[0097] For the particles updated by the two-level correlation filter new weights are assigned and the weights are normalized. Let the particle set updated by the two-level correlation filter be

[0098] Further, regarding step 6, on the premise of the target centroid dynamics estimation state of the current frame image obtained in step 5, the target color feature, target edge feature, and the corresponding target color feature correlation filter and target edge feature correlation filter of the current frame are updated to obtain the updated target color feature, target edge feature, and the corresponding target color feature correlation filter and target edge feature correlation filter of the current frame. The specific content is as follows:

[0099] Centered on the position of the target centroid dynamics estimation state of the current frame image obtained in step 5, a window with a size 2.5 times that of the target bounding box is selected for the current frame image. The cyclic shift method is used for the selected window to construct a training sample set, and the color feature and edge feature of the training sample set are respectively extracted to obtain the target color feature of the current frame image and the target edge feature of the current frame image The ridge regression method is used to train the color feature and edge feature of the current frame image respectively to obtain the target color feature correlation filter of the current frame image and the target edge feature correlation filter

[0100] Based on the target color feature correlation filter of the previous frame image the target edge feature correlation filter and the target color feature correlation filter of the current frame image and the target edge feature correlation filter as the premise, the target color feature correlation filter and target edge feature correlation filter of the current frame image are respectively updated to obtain the updated target color feature correlation filter of the current frame image and the target edge feature correlation filter The update formula adopted is as follows:

[0101]

[0102]

[0103] Similarly, based on the target color feature of the previous frame image the target edge feature and the obtained target color feature of the current frame image above the target edge feature the target color feature and target edge feature of the current frame image are respectively updated to obtain the updated target color feature of the current frame image and the target edge feature The update formula adopted is as follows:

[0104]

[0105]

[0106] Where η is the learning factor, and η ∈ [0, 1].

[0107] In this embodiment, the device operating environment is Windows 10, x64-bit operating system, the processor is Inter(R) core(TM) i7-8700 CPU (3.2GHz), and the RAM is 8GB. The instance simulation experiment is implemented using MATLAB R2018a coding. In this instance, the number of particles is set to N = 5, the regularization parameter is λ = 0.0001, and the learning factor for template update is η = 0.99.

[0108] To further verify the effectiveness of the related particle filtering method for multi-feature cascade fusion provided by the present invention, the publicly available dataset OTB2013 and 50 groups of video sequences are used to test the tracking effect. Among the 50 groups of video image sequences, there are image sequences with different attributes, such as target tracking tests in complex environments such as illumination change, scale change, and target occlusion. To illustrate the superiority of the method of the present invention, the present invention is compared and analyzed with existing visual tracking methods Struck, TLD method, particle filtering method SCM, and correlation filtering methods KCF and CSK.

[0109] The tracking performance of the algorithm adopted in the present invention is evaluated and analyzed in terms of tracking success rate and tracking accuracy in three modes: OPE (one-pass evaluation), TRE (temporal robustness evaluation), and SRE (spatial robustness evaluation).

[0110] Figure 2 shows the analysis of the tracking performance of the algorithm with the tracking success rate of the target as the evaluation index. The success rate of single-frame image tracking is judged by the overlap rate. The calculation formula of the overlap rate is

[0111]

[0112] Where A is the area of the tracking result region and B is the area of the true tracking region.

[0113] As can be seen from Figure 2, the tracking performance of the method of the present invention is superior to other tracking algorithms in all three modes.

[0114] Figure 3 shows the algorithm analyzing the tracking performance with the tracking accuracy of the target as the evaluation index. The tracking accuracy is calculated by the ratio of the number of frames in which the Euclidean distance between the true centroid position and the estimated centroid position of the target bounding box is less than the threshold to the total number of frames. Figure 3 respectively shows the tracking accuracy curves of different algorithms under different thresholds in three modes.

[0115] The method adopted by the present invention can well handle the tracking of video sequences under conditions such as illumination changes, scale changes, and target occlusion. To further illustrate the tracking advantages of the algorithm under these three attributes, Figure 4 gives the tracking success rates of the listed tracking algorithms in the OPE mode. It can be seen from the figure that the method of the present invention has strong robustness for the tracking of video sequences under complex backgrounds such as illumination changes, scale changes, and target occlusion.

Claims

1. A related particle filter video tracking method based on multi-feature cascade fusion, characterized in that: The method includes the following steps: Step 1: Obtain a set of video images containing M frames, and perform initialization settings on the initial frame image, including initializing the centroid dynamics state of the target in the initial frame image, initializing the centroid particle set, and initializing the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image; Step 2: Read the frame image corresponding to the next frame, and predict the particle motion state of the current frame image according to the dynamic motion state of the particles corresponding to the previous frame image, so as to obtain the predicted state of the particle set corresponding to the current frame image; Step 3: Perform a first-level update on the predicted state of the particle set of the current frame image according to the target color feature correlation filter corresponding to the previous frame image; Step 4: Perform a second-level update on the estimated state of the particle set of the current frame image obtained in Step 3 according to the target edge feature correlation filter corresponding to the previous frame image; Step 5: Process the estimated state of the particle set of the current frame image obtained in Step 4 to obtain the centroid dynamics estimation state of the target in the current frame image; Step 6: Based on the centroid dynamics estimation state of the target in the current frame image obtained in Step 5, update the target color feature, the target edge feature, and the corresponding target color feature correlation filter and target edge feature correlation filter of the current frame image; The Step 6 updates the target color feature correlation filter and the target edge feature correlation filter corresponding to the current frame image based on the centroid dynamics estimation state of the target of the current frame image obtained in Step 5, and is implemented by the following formula: Wherein: is the relevant filter for the target color feature of the updated current frame (the k-th frame), is the relevant filter for the target color feature of the previous frame (the (k - 1)-th frame) image, is the relevant filter for the target color feature of the current frame (the k-th frame) before update; is the relevant filter for the target edge feature of the updated current frame (the k-th frame), is the relevant filter for the target edge feature of the previous frame (the (k - 1)-th frame) image, is the relevant filter for the target edge feature of the current frame (the k-th frame) before update, η k is the learning factor η k ∈[0, 1]; Step 7: Repeat Step 2 - Step 6 until the processing of each frame image in the M-frame video image is completed and then end.

2. The multi-feature cascade fusion related particle filter video tracking method according to claim 1, wherein: The Step 1 initializes the dynamics state of the centroid of the target in the initial frame image, where the dynamics equation of the target centroid is: x k = Ax k-1 + ω k where A is the state transition matrix of the target centroid motion, T is the sampling time interval, and ω k is Gaussian white noise with zero mean; set the initial dynamic state of the target centroid as x0 and y0 represent the abscissa and ordinate of the target centroid particle respectively, represent the moving speeds of the target centroid particle in the x-direction and y-direction respectively.

3. The multi-feature cascaded fusion related particle filter video tracking method according to claim 1, wherein: The process of initializing the target centroid particles in the Step 1 is as follows: Set the target bounding box centered on the target centroid, use the initial dynamics state of the target centroid as the mean of the normal distribution, construct the variance of the normal distribution with the diagonal of the target bounding box, perform random sampling on the target bounding box to obtain N random particles, and assign equal weights to the N particles.

4. The related particle filter video tracking method with multi-feature cascade fusion according to claim 1, characterized in that: The process of initializing the target color feature correlation filter and the target edge feature correlation filter corresponding to the initial frame image in the Step 1 is as follows: Centered on the centroid of the target bounding box, select a window with a size 2.5 times that of the target bounding box, construct a training sample set for the window using the cyclic shift method, extract the color features and edge features of the training sample set, and use the ridge regression method to train the color features and edge features respectively to obtain the target color feature correlation filter and the target edge feature correlation filter.

5. The multi-feature cascaded fusion related particle filter video tracking method according to claim 1, wherein: The specific process of the Step 3 performing a first-level update on the predicted state of the particle set of the current frame image according to the target color feature correlation filter corresponding to the previous frame image is as follows: Step 3.1: Centering on the position of each particle in the current frame image, a search region with a size 2.5 times that of the target bounding box is selected. The cyclic shift method is used to construct a sample set for this search region, and the color features of the sample set are extracted to obtain the color features of each particle in the current frame image; Step 3.2: For the color features corresponding to the current frame image, the color features of the particles in the current frame image are successively subjected to window function processing and fast Fourier transform to obtain the processed color features. The processed color features are multiplied by the color feature correlation filter corresponding to the previous frame image, and the multiplication result is subjected to inverse fast Fourier transform to obtain the target color feature correlation response map of each particle in the current frame image; Step 3.3: The positions of the maximum response values of the target color feature correlation response maps of each particle in the current frame image obtained in Step 3.2 are respectively used as the new positions of the corresponding particles in the current frame image, and the maximum response values of the target color feature correlation response maps of each particle are used as the weight values of the corresponding particles in the current frame image. Then, the obtained particle weight values are normalized to obtain the estimated state of the particle set in the current frame image after the first-level update.

6. The related particle filter video tracking method with multi-feature cascade fusion according to claim 1, wherein: The following is the specific process of Step 4 for performing a second-level update on the estimated state of the particle set of the current frame image obtained in Step 3 according to the target edge feature correlation filter corresponding to the previous frame image: Step 4.1: Centering on the position of each particle in the current frame image obtained in Step 3, a search region with a size 2.5 times that of the target bounding box is selected. The cyclic shift method is used to construct a sample set for this search region, and the edge features of the sample set are extracted to obtain the edge features of each particle in the current frame image; Step 4.2: For the edge features corresponding to the current frame image, the edge features of the particles in the current frame image are successively subjected to window function processing and fast Fourier transform to obtain the processed edge features. The processed edge features are multiplied by the edge feature correlation filter corresponding to the previous frame image, and the multiplication result is subjected to inverse fast Fourier transform to obtain the target edge feature correlation response map of each particle in the current frame image; Step 4.3: The positions of the maximum response values of the target edge feature correlation response maps of each particle in the current frame image obtained in Step 4.2 are respectively used as the positions of the corresponding particles in the current frame image, and the maximum response values of the target edge feature correlation response maps are used as the weight values of the corresponding particles in the current frame image. Then, the obtained particle weight values are normalized to obtain the estimated state of the particle set in the current frame image after the second-level update.

7. The multi-feature cascaded fusion related particle filter video tracking method according to claim 1, characterized in that: Based on the premise of the target centroid dynamics estimated state of the current frame image obtained in Step 5, the target color features and target edge features corresponding to the current frame image are updated, which is achieved by using the following formula: Among them, the target color feature corresponding to the updated current frame (the k-th frame) image, the target color feature of the previous frame (the (k - 1)-th frame) image, the target color feature corresponding to the current frame (the k-th frame) image before update, the target edge feature corresponding to the updated current frame (the k-th frame) image, the target edge feature of the previous frame (the (k - 1)-th frame) image, the target edge feature corresponding to the current frame (the k-th frame) image before update, η k is the learning factor η k ∈[0, 1].