Anti-occlusion Moving Target Tracking Device and Method Based on ROI Prediction and Multi-module Learning
By adopting ROI prediction and multi-module learning methods in mobile target tracking systems, the problems such as target occlusion, scale changes and lighting changes in long-term tracking are solved, and the robustness and real-timeness of the system are improved, and the secondary tracking with high accuracy is achieved.
Patent Information
- Application Number
- CN202210615297.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The prior art has problems such as target occlusion, target disappearance, motion mutation, scale changes, and lighting changes in the long-term tracking of a single moving target, resulting in low robustness, poor real-timeness and large calculation amount.
An anti-occlusion mobile target tracking device based on ROI prediction and multi-module learning is adopted, including a tracking module, a detection module, a learning module and a comprehensive module. The tracking module performs position and scale prediction through multi-feature extraction and related filters, the detection module uses ROI prediction and fHOG-SVM classifier for object detection, the learning module optimizes the classifier performance through P-N learning, and the comprehensive module coordinates the multi-module output.
It improves the robustness, accuracy and real-time nature of target tracking, can effectively respond to the scale changes, lighting changes and occlusion of targets, realizes secondary tracking, and reduces the amount of calculation and resource waste.
Smart Images

Figure CN114972735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual target detection, in particular to an anti-occlusion moving target tracking device and method based on ROI prediction and multi-module learning. Background Art
[0002] Visual target tracking is a challenging computer vision task and one of the key technologies in the field of artificial intelligence. Visual target tracking is defined as processing each frame of the collected video, and the processing process includes detecting, positioning, identifying, and tracking the target, so as to dynamically obtain the position information and motion information of the target. Visual target tracking is an indispensable part in fields such as mobile navigation, video surveillance, and human-computer interaction, and has good application prospects and development space.
[0003] Although some progress has been made in the field of target tracking, for the long-term tracking of a single moving target, there are still problems such as target occlusion, target disappearance, motion mutation, scale change, and illumination change. Currently, many algorithms cannot handle the long-term tracking problem. This is because during the long-term movement of the target, it is easier to be lost and its appearance changes, so a secondary detection mechanism needs to be integrated. To solve the long-term tracking problem, the framework of the tracking-detection-learning algorithm is used as the basic framework. Tracking-detection-learning is a robust tracking algorithm. The tracking module tracks the target when the target is continuously visible. The learning module can continuously learn the appearance of the target and update the target sample library in real time. The detection module performs real-time detection of the target globally and can still detect the target when the target is occluded or disappears and then reappears. The framework of multi-module learning and collaborative update of the tracking-detection-learning algorithm is the key to ensuring long-term tracking. However, for many current algorithms using the framework of the tracking-detection-learning algorithm, the tracking module has low robustness when encountering problems such as scale change and illumination change. Under the limitations of the embedded platform, the algorithm has a large amount of calculation and low real-time performance. In addition, the detection module has many sliding windows, wasting resources and slowing down the detection speed. Therefore, by studying the tracking-detection-learning framework and improving two important modules, improving the robustness, accuracy, and real-time performance of the algorithm and realizing secondary detection are the key points to solve the long-term tracking problem of moving targets. Summary of the Invention
[0004] The purpose of the present invention is to provide an anti-occlusion moving target tracking device and method based on ROI prediction and multi-module learning with high accuracy, high real-time performance, and high robustness.
[0005] The technical solution for realizing the purpose of the present invention is: an anti-occlusion moving target tracking device based on ROI prediction and multi-module learning, including a tracking module, a detection module, a learning module, and a comprehensive module, wherein:
[0006] The tracking module includes a multi-feature extraction module and a correlation filtering module. Based on the position filter, a scale filter is added as the algorithm framework of the tracking module. By adding a multi-feature extraction and fusion optimization framework, PCA dimensionality reduction and QR decomposition are performed on the position filter and the scale filter;
[0007] The detection module includes an ROI prediction module and a cascade classification module. The ROI prediction module uses square root cubature Kalman filtering for position estimation, obtains the ROI region using the estimated position and uses it as the input of the cascade classification module. The fHOG-SVM classifier is used in the cascade classification module;
[0008] The tracking module and the detection module work synchronously, mutually correct, learn, and update parameters through the learning module. The comprehensive module obtains the final output position through the coordinated work among multiple modules, realizing the tracking of a single moving target.
[0009] An occlusion-resistant moving target tracking method based on ROI prediction and multi-module learning, comprising the following steps:
[0010] Step 1, multi-feature extraction: Convert the input RGB three-channel image into a grayscale single-channel image, standardize the color space of the image using the gamma correction method, calculate the image gradient, including the gradient value and gradient direction of each pixel point, and then construct a 9-dimensional HOG feature vector; Obtain a 36-dimensional feature vector corresponding to each cell through normalization truncation, perform PCA dimensionality reduction, extract 31-dimensional features, combine the features of each cell, and obtain an M×N×31-dimensional fHOG feature from an M×N image, and splice it with an M×N×1 grayscale feature to obtain an M×N×32-dimensional fusion feature;
[0011] Step 2, correlation filtering tracking: Set a position filter and a scale filter. For the position filter, first initialize the expected two-dimensional Gaussian output of the target, collect a sample centered on the target position, reduce the fusion feature from 32 dimensions to 18 dimensions using PCA dimensionality reduction, extract these 18-dimensional features for each pixel point of the sample, multiply by a two-dimensional Hamming window, and use it as the test input, and then use the inverse Fourier transform to determine the new position of the target; For the scale filter, first initialize the expected one-dimensional Gaussian output of the scale filter, extract samples at different scales centered on the target position, pass each sample through a one-dimensional Hamming window, use it as the test input, and then use the inverse Fourier transform to determine the new scale of the target;
[0012] Step 3, ROI prediction area: Use the position of the target in the image at the previous moment as the observation value, and use the square root cubature Kalman filter algorithm to estimate the position of the target in the image at the next moment. Define a region based on the aspect ratio of the previous frame and four times the area, and send this region as the ROI region of the current frame to the detection module;
[0013] Step 4, cascade classification: Set up an image meta-variance classifier, an fHOG-SVM classifier, and a nearest neighbor classifier. The ROI region is the area where the target in the current frame is most likely to appear. Use this ROI region as the input to the cascade classification module, that is, the area to be detected. First, generate multiple sliding windows of different scales in the area to be detected and send them to the image meta-variance classifier. Calculate the pixel gray variance between the test window and the target box image. Consider the test samples with a variance less than half of the target sample variance as negative samples. Then, use the positive samples as the input to the fHOG-SVM classifier, extract the fHOG features, and send them to the SVM classifier to obtain the positive and negative sample classification results. Finally, use the positive sample windows obtained from the first two classifiers as the input to the nearest neighbor classifier, match the similarity of each window with the online model in turn, and update the positive sample space of the online model to obtain the final positive samples of the cascade classification module, which are the outputs of the detection module;
[0014] Step 5, learning and updating: Use the P-N learning method to optimize the performance of the classifier in the detection module in an online learning manner. In P-N learning, first, use the tracking module to predict the position of the target in the current frame. If the predicted position is detected as a negative sample by the detection module, the P expert will correct the positive sample that was misclassified as a negative sample to a positive sample and send it to the training set. Then, the N expert compares the positive samples generated by the detection module with the positive samples obtained by the P expert, selects the most credible samples as the output position;
[0015] Step 6, multi-module integration: Through the coordinated work of multiple modules, obtain the final output position to achieve the tracking of a single moving target.
[0016] Compared with the prior art, the significant advantages of the present invention are as follows: (1) On the basis of the position filter, a scale filter is added as the algorithm framework of the tracking module. By adding multi-feature extraction and fusion to improve and optimize the framework, the tracking robustness under illumination changes is enhanced. PCA dimensionality reduction and QR decomposition are performed on the position filter and the scale filter to reduce the computational amount and improve the real-time performance; (2) In the detection module, square root cubature Kalman filtering is used for position prediction, which improves the detection robustness when the target is occluded. The ROI region is obtained using the predicted position and used as the input of the cascade classifier, narrowing the search range, reducing the computational amount, and improving the real-time performance; (3) In the cascade classifier, the fHOG-SVM classifier is used instead of the original random fern classifier. The fHOG feature not only maintains good invariance to both image scale and illumination changes, but also reduces the computational amount compared with the HOG feature, improving the algorithm speed. The combination of fHOG and SVM improves the detection accuracy and real-time performance; (4) Although the tracking module and the detection module work independently inside the module, they learn from each other and update parameters through the learning and updating module, and jointly confirm the final predicted position through the comprehensive module. When the target reappears after the occlusion disappears and the tracking module cannot work properly, the detection module can still detect the target through initialization detection and re-initialize the tracking module with the result, and the module restarts to complete secondary tracking. Description of the Drawings
[0017] Figure 1 It is a framework diagram of the anti-occlusion moving target tracking method based on ROI prediction and multi-module learning of the present invention. Detailed Embodiment
[0018] An anti-occlusion moving target tracking device based on ROI prediction and multi-module learning of the present invention includes a tracking module, a detection module, a learning module, and a comprehensive module, wherein:
[0019] The tracking module includes a multi-feature extraction module and a correlation filtering module. On the basis of the position filter, a scale filter is added as the algorithm framework of the tracking module. By adding multi-feature extraction and fusion to optimize the framework, PCA dimensionality reduction and QR decomposition are performed on the position filter and the scale filter;
[0020] The detection module includes an ROI prediction module and a cascade classification module. The ROI prediction module uses square root cubature Kalman filtering for position prediction, obtains the ROI region using the predicted position and uses it as the input of the cascade classification module, and the fHOG-SVM classifier is used in the cascade classification module;
[0021] The tracking module and the detection module work synchronously, correct, learn, and update parameters through the learning module, and the comprehensive module obtains the final output position through the coordinated work among multiple modules to realize the tracking of a single moving target.
[0022] As a specific example, the multi-feature extraction module is obtained by splicing and fusing the grayscale feature of M×N×1 and the fast histogram of oriented gradients (fHOG) feature of M×N×31. Here, both M and N are positive integers. The 31-dimensional fHOG feature is obtained by first normalizing and truncating to get the 36-dimensional feature vector corresponding to each cell and then using principal component analysis (PCA) for dimensionality reduction, including 18-dimensional signed fHOG gradients, 9-dimensional unsigned fHOG gradients, and 4-dimensional features from the normalization operations of 4 neighboring cells of the current cell and the diagonal neighboring cells.
[0023] As a specific example, the basic framework of the correlation filtering module is the discriminative scale space tracking algorithm framework. The position filter and the scale filter are used to perform target localization and scale evaluation in sequence, and principal component analysis (PCA) dimensionality reduction and orthogonal triangular (QR) decomposition are respectively performed on the position filter and the scale filter to optimize the algorithm framework.
[0024] As a specific example, the ROI prediction module estimates the ROI region by adding square root cubature Kalman filtering; with the target position predicted in the current frame as the center, a region is delimited based on the aspect ratio of the previous frame and four times the area, and this region is used as the ROI region of the current frame and sent to the cascade classification module.
[0025] As a specific example, the cascade classification module includes an image meta-variance classifier, an fHOG-SVM classifier, and a nearest neighbor classifier, specifically as follows:
[0026] The ROI prediction module predicts the region where the target in the current frame is most likely to appear, and uses this as the input of the cascade classification module, that is, the region to be detected. First, multiple test sliding windows of different scales are generated in the region to be detected and sent to the image meta-variance classifier. The pixel grayscale variance between the test window and the target box image is calculated, and the test samples with a variance less than half of the target sample variance are considered negative samples; then the positive samples are used as the input of the fHOG-SVM classifier, the fHOG features are extracted and sent to the SVM classifier to obtain the positive and negative sample category results; finally, the positive sample windows obtained by the first two classifiers are used as the input of the nearest neighbor classifier, the similarity between each window and the online model is matched in sequence, and the positive sample space of the online model is updated, so as to obtain the final positive samples of the cascade classification module, that is, the output of the detection module.
[0027] A method for tracking occluded moving targets based on ROI prediction and multi-module learning according to the present invention includes the following steps:
[0028] Step 1, Multi-feature extraction: Convert the input RGB three-channel image to a grayscale single-channel image, standardize the color space of the image using the gamma correction method, calculate the image gradient, including the gradient value and gradient direction of each pixel point, and then construct a 9-dimensional HOG feature vector; obtain a 36-dimensional feature vector corresponding to each cell through normalization truncation, reduce the dimension through PCA to extract 31-dimensional features, combine the features of each cell, and obtain an M×N×31-dimensional fHOG feature from an M×N image, and splice it with the M×N×1 grayscale feature to obtain an M×N×32-dimensional fusion feature;
[0029] Step 2, Correlation filter tracking: Set up a position filter and a scale filter. For the position filter, first initialize the expected two-dimensional Gaussian output of the target, collect a sample centered on the target position, reduce the dimension of the fusion feature from 32 to 18 using PCA, extract these 18-dimensional features for each pixel point of the sample, multiply by a two-dimensional Hamming window, and use it as the test input, and then use the inverse Fourier transform to determine the new position of the target; for the scale filter, first initialize the expected one-dimensional Gaussian output of the scale filter, extract samples at different scales centered on the target position, pass each sample through a one-dimensional Hamming window, use it as the test input, and then use the inverse Fourier transform to determine the new scale of the target;
[0030] Step 3, ROI prediction region: Use the position of the target in the image at the previous moment as the observation value, estimate the position of the target in the image at the next moment using the square root cubature Kalman filter algorithm, delimit the region with the aspect ratio of the previous frame and four times the area, and send this region as the current frame ROI region into the detection module;
[0031] Step 4, Cascade classification: Set up an image meta-variance classifier, an fHOG-SVM classifier, and a nearest neighbor classifier. The ROI region is the region where the target in the current frame is most likely to appear. Use this ROI region as the input of the cascade classification module, that is, the region to be detected. First, generate multiple different-scale detection sliding windows in the region to be detected and send them into the image meta-variance classifier, calculate the pixel gray variance between the detection window and the target box image, and consider the test samples with a variance less than half of the target sample variance as negative samples; then use the positive samples as the input of the fHOG-SVM classifier, extract the fHOG features, and send them into the SVM classifier to obtain the positive and negative sample class results; finally, use the positive sample windows obtained by the first two classifiers as the input of the nearest neighbor classifier, match the similarity between each window and the online model in turn, and update the positive sample space of the online model, so as to obtain the final positive samples of the cascade classification module, that is, the output of the detection module;
[0032] Step 5, Learning and Updating: Use the P-N learning method to optimize the performance of the classifier in the detection module in an online learning manner; in P-N learning, first, use the tracking module to predict the target position in the current frame. If the predicted position is detected as a negative sample by the detection module, the P expert will correct the positive sample that has been misclassified as a negative sample to a positive sample and send it to the training set; then, the N expert will compare the positive samples generated by the detection module with the positive samples obtained by the P expert and select the most credible samples as the output positions.
[0033] Step 6, Multi-module Integration: Through the coordinated work of multiple modules, obtain the final output position to achieve the tracking of a single moving target.
[0034] As a specific example, in Step 2, for the position filter, first initialize the expected two-dimensional Gaussian output of the target, collect a sample centered on the target position, use PCA dimensionality reduction to reduce the fused features from 32 dimensions to 18 dimensions, extract these 18-dimensional features for each pixel point of the sample, multiply by a two-dimensional Hamming window as the test input, and then use the inverse Fourier transform to determine the new position of the target, as follows:
[0035] First, let the selected target sample in the initial image be the positive sample f, and select the two-dimensional Gaussian function as the expected output sample g, so that the following formula is minimized:
[0036]
[0037] where, * represents the convolution operation, λ represents the weight coefficient, f l represents the feature of the l-th channel, h l represents the filter of the l-th channel, l ∈ {1, 2,..., d}, and d is the dimension of the selected features;
[0038] Convert the above formula to the complex frequency domain and solve it using Parseval's formula:
[0039]
[0040] where, H l , F l , G are the corresponding variables obtained by performing the discrete Fourier transform DFT on h l , f l , g, is the conjugate transpose of G;
[0041] Use a training sample f t to update the filter parameters:
[0042]
[0043]
[0044] Among them, and B t are the numerator and denominator of a training sample f corresponding to the filter t respectively, and B t-1 are the numerator and denominator of the filter of the previous frame of this training sample, and η is the learning rate;
[0045] If z t is an image sample, is the variable obtained by discrete Fourier transform, then the output y t is:
[0046]
[0047] Among them, and are the numerator and denominator of the filter in the previous frame; y t is the correlation score. By finding the maximum correlation score, the state estimation of the current target position is obtained.
[0048] As a specific example, in step 3, the position of the target in the image at the previous moment is used as the observation value, and the square root cubature Kalman filtering algorithm is used to estimate the position of the target in the image at the next moment. The aspect ratio of the previous frame and four times the area are used to delimit the region, and this region is used as the ROI region of the current frame and sent to the detection module, specifically as follows:
[0049] For a discrete nonlinear dynamic target tracking system with additive noise:
[0050]
[0051] Among them, x t and y t represent the state and measurement value of the system at time t respectively, f(·) and h(·) are the nonlinear state transition function and nonlinear measurement function respectively, the process noise w t-1 and the measurement noise v t-1 are independent of each other, and w t-1 ~N(0, Q t ), v t-1 ~N(0, R t );
[0052] State estimation includes time update and measurement update. When a fault is detected, the state parameters x t-1 and S t-1 of the last successful frame are used to initialize the filter, and then the filter gain K t and the new state estimation The square root factor S of the sum of errors and covariance t :
[0053]
[0054] where P xy,t is the cross-covariance matrix of the measurement prediction values, S yy,t is the square root of the auto-covariance matrix of the measurement prediction values, is the system state prediction value, is the estimated measurement state prediction value, χ t , γ t is the weight matrix, S R,t is the square root of the covariance matrix of the measurement noise;
[0055] Taking the position v = (i, j) of the target in the image at the previous moment as the observation value, estimating the position of the target in the image at the next moment Using the aspect ratio of the previous frame and quadruple the area to delimit the region, and sending this region as the current frame ROI region into the detection module.
[0056] As a specific example, in step 4, SVM uses a kernel function to solve the non-linear problem. By establishing a hyperplane in the feature space as the decision surface, the separation margin between the positive samples and the negative samples is maximized to separate the positive and negative samples. Assuming the hyperplane is:
[0057] wx + b = 0
[0058] where w represents the normal vector, determining the direction of the hyperplane, and b represents the offset, determining the distance between the hyperplane and the origin;
[0059] The training sample set train = {(x 1 , y 1 ), (x 2 , y 2 ),..., (x n , y n )}, x ∈ R n , y i ∈ {+1, -1}, i represents the i-th sample, and n represents the sample size; the classification surface needs to satisfy y i [wx i + b] ≥ 1, i = 1, 2,..., m, and the optimal hyperplane problem is transformed into:
[0060]
[0061] Introducing the Lagrangian function:
[0062]
[0063] should satisfy and the optimal solution is obtained the optimal weight method vector w * , the optimal offset b * :
[0064]
[0065]
[0066] So the optimal hyperplane w * x + b * = 0, and the optimal classification function f(x) = sgn{w * x + b *};
[0067] Compare the positive samples output by the fHOG - SVM classifier with the online model for similarity, implement sample classification, and update the positive sample space of the online model. The similarity calculation is as follows:
[0068]
[0069] Among them, S r is the correlation similarity, S + is the positive similarity, S - is the negative similarity, and the definitions are as follows:
[0070]
[0071]
[0072] Among them, M represents the target model of the sample library, represents the positive sample, represents the negative sample, and p represents the sample to be measured;
[0073] The calculation formula of S is as follows:
[0074] S(p i , p j ) = 0.5(NCC(p i , p j ) + 1)
[0075] Among them, the definition of NCC is as follows:
[0076]
[0077] Among them, μ i , σ i are the mean and standard deviation of the image patch p i , μ j , σ j are the mean and standard deviation of the image patch pj The mean and standard deviation of;
[0078] Finally, the calculated S r is compared. The larger S r is, the greater the likelihood that the sample is the target. Set a threshold γ. Samples with S r > γ are considered positive samples, and vice versa are negative samples and discarded; at the same time, the new positive samples are added to the positive sample library of the online model for subsequent matching. The number of the positive sample library of the online model is fixed. If the number is insufficient, add some samples. If it exceeds the upper limit of the number, randomly delete some samples and then add new samples.
[0079] As a specific example, step 6 is specifically as follows:
[0080] The comprehensive module obtains the final output position through the coordinated work among multiple modules. According to the operation results of the detection module and the tracking module, it is divided into four cooperation modes:
[0081] (1) Tracking success and detection success
[0082] Detection success means that at least one sliding window passes through the detection module, and after clustering the passed sliding windows, the final clustering result has only one clustering center. Detection failure means that no sliding window passes through the detection module, or multiple sliding windows pass through the detection module but the clustering result has multiple clustering centers; Tracking success means that the tracking module outputs a feature rectangle box, and tracking failure means that no feature rectangle box is output;
[0083] If tracking is successful and detection is successful, cluster the detection results to obtain relevant output results, and judge the overlap rate and credibility between the clustering center and the tracking module. If the overlap degree between the two is lower than the threshold of 0.5 and the credibility of the detection module is high, use the detection module to correct the result of the tracking module. If the overlap degree is higher than the threshold of 0.5, use the weighted average of the detection module and the tracking module to obtain the result as the final output;
[0084] (2) Tracking success and detection failure
[0085] If tracking is successful but detection fails, directly use the output of the tracking module as the final output of the current frame;
[0086] (3) Tracking failure and detection success
[0087] If the tracking fails but the detection is successful, cluster the output sample boxes of the detection module. If there is only one clustering center in the final clustering result, use the result of this clustering as the final output and re-initialize the tracking module with the result of this clustering, that is, the re-detection process that enters again after the target disappears; if there are multiple clustering centers, it means that although the detection module has passed, there are multiple different positions, and it is considered that the detection has failed.
[0088] (4) Tracking fails and detection fails
[0089] If both the tracking module and the detection module fail, it is considered that this detection is invalid and discarded.
[0090] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0091] Embodiment
[0092] Combined with Figure 1 , the anti-occlusion moving target tracking method of the present invention based on ROI prediction and multi-module learning includes a feature extraction module, a correlation filtering tracking module, an ROI prediction module, a cascaded detection module, a learning module, and a comprehensive module. The specific algorithm real-time process is as follows:
[0093] Step 1: Multi-feature extraction
[0094] Convert the input RGB three-channel image into a single-channel image by graying, standardize the color space of the image using the gamma correction method, calculate the image gradient, including the gradient value and gradient direction of each pixel point, and then construct a 9-dimensional HOG feature vector. Obtain a 36-dimensional feature vector corresponding to each cell through normalization truncation. Extract 31-dimensional features through PCA dimensionality reduction, combine the features of each cell, and obtain M×N×31-dimensional fHOG features from an M×N image, and splice them with M×N×1 gray features to obtain M×N×32-dimensional fused features. Through multi-feature fusion, the robustness to rapid changes in illumination and shape is improved.
[0095] Step 2: Correlation filtering tracking
[0096] The correlation filtering tracking module includes two parts, a position filter and a scale filter. First, initialize the expected two-dimensional Gaussian output of the target, collect a sample centered on the target position. To reduce the computational amount and improve the running speed, use PCA dimensionality reduction to reduce the fused features from 32 dimensions to 18 dimensions, extract these 18-dimensional features for each pixel point of the sample, multiply by a two-dimensional Hamming window, and use it as the test input, and then use the inverse Fourier transform to obtain y t, the maximum value of y is the target new position. The scale filter adopts the same design method as the position filter. First, initialize the expected one-dimensional Gaussian output of the scale filter. Centered on the target position, extract samples at different scales, pass each sample through a one-dimensional Hamming window as the test input, and then use the inverse Fourier transform to obtain y t , the maximum value of y is the target new scale.
[0097] Step 3: ROI prediction region
[0098] Since the detection module needs to generate multi-scale sliding windows as samples in the global image range, which greatly increases the computational complexity and reduces the algorithm speed. Therefore, in the present invention, by predicting the target position at the next moment, the ROI region, i.e., the region of interest, is determined to narrow the search range, reduce the detection samples, and improve the real-time performance. Take the position v=(i,j) of the target in the image at the previous moment as the observation value, and use the square root cubature Kalman filter algorithm to estimate the position v=(i,j) of the target in the image at the next moment. Define a region based on the aspect ratio of the previous frame and four times the area, and send this region as the current frame ROI region into the detection module.
[0099] Step 4: Cascade detection
[0100] The cascade detection module includes an image meta-variance classifier, an fHOG-SVM classifier, and a nearest neighbor classifier.
[0101] The ROI prediction module predicts the region where the target in the current frame is most likely to appear, and uses this as the input of the cascade classifier, i.e., the region to be detected. First, obtain the samples to be detected by using the multi-scale displacement method in the region to be detected, and send them into the image meta-variance classifier. Calculate the pixel gray variance between the window to be detected and the target box image respectively. If the variance of the sample to be detected is less than half of the variance of the target box image, it is considered a negative sample. By variance screening, half of the scanning frames in the input region can be reduced.
[0102] Then, the positive samples obtained by the image meta-variance classifier are used as the input of the fHOG-SVM classifier. Extract the fHOG features and send them into the SVM classifier to obtain the classification results of positive and negative samples. SVM uses a kernel function to solve the non-linear problem. Its main idea is to establish a hyperplane in the feature space as the decision surface, so that the separation margin between positive and negative samples is maximized to separate positive and negative samples.
[0103] Compare the similarity between the positive samples output by the fHOG-SVM classifier and the online model to achieve sample classification and update the positive sample space of the online model. The similarity calculation is as follows:
[0104]
[0105] where Sr is the relevant similarity, S + is the positive similarity, S - is the negative similarity, defined as follows:
[0106]
[0107]
[0108] where M represents the target model of the sample library, represents the positive sample, represents the negative sample, and p represents the sample to be tested. The calculation formula of S is as follows:
[0109] S(pi, pj) = 0.5(NCC(pi, pj) + 1)
[0110] where NCC is defined as follows:
[0111]
[0112] where μ i , σ i are the mean and standard deviation of the image patch p i , μ j , σ j are the mean and standard deviation of the image patch p j .
[0113] Finally, compare the calculated S r , the larger S r , the greater the possibility that the sample is the target. Set the threshold γ, and the sample with S r > γ is considered a positive sample, otherwise it is a negative sample and is discarded. At the same time, add the new positive samples to the positive sample library of the online model for subsequent matching. The number of the positive sample library of the online model is fixed. If the number is insufficient, add samples. If it exceeds the upper limit of the number, randomly delete some samples and then add new samples.
[0114] Step Five: Learning and Updating
[0115] The P-N learning method is used in the algorithm to optimize the performance of the classifier in the detection module in an online learning manner and improve the generalization ability of the classifier. In P-N learning, first, the tracking module is used to predict the target position of the current frame. If the predicted position is detected as a negative sample by the detection module, the P expert will correct the positive sample that is misclassified as a negative sample to a positive sample and send it to the training set. Then, the N expert will compare the positive samples generated by the detection module with the positive samples obtained by the P expert and select the most credible sample as the output position.
[0116] Step Six: Multi-module Integration
[0117] The comprehensive module obtains the final output position through the coordinated work of multiple modules. According to the running results of the detection module and the tracking module, it can be divided into four cooperation modes:
[0118] (1) Tracking successful, detection successful
[0119] If the tracking is successful and the detection is successful, cluster the detection results to obtain the relevant output results, and judge the clustering center, the overlap rate and the credibility with the tracking module. If the overlap between the two is low and the credibility of the detection module is high, use the detection module to correct the results of the tracking module. If the overlap is close, use the weighted average of the detection module and the tracking module to obtain the result as the final output.
[0120] (2) Tracking successful, detection failed
[0121] If the tracking is successful but the detection fails, directly use the output of the tracking module as the final output of the current frame.
[0122] (3) Tracking failed, detection successful
[0123] If the tracking fails but the detection is successful, cluster the output sample frames of the detection module. If there is only one clustering center in the final clustering result, use this result as the final output and re-initialize the tracking module with this. This is also the re-detection process when the target reappears after disappearing. If there are multiple clustering centers, it means that although it has passed the detection module, there are multiple different positions, and it is also considered that the detection has failed.
[0124] (4) Tracking failed, detection failed
[0125] If both the tracking and detection modules fail, it is considered that this detection is invalid and discarded.
[0126] In view of the problems such as scale change, illumination change, target occlusion, and target disappearance during the long-term tracking of a single moving target, the present invention provides a target tracking method, which can not only overcome the above problems to achieve secondary tracking, but also has high real-time performance and high robustness. The present invention also provides an anti-occlusion moving target tracking device based on ROI prediction and multi-module learning, which includes four parts: a tracking module, a detection module, a learning module, and a comprehensive module. The tracking module based on multi-feature extraction and correlation filtering algorithm and the detection module based on ROI prediction and fHOG-SVM are used to estimate the target position of each single frame respectively, and the comprehensive module is used to comprehensively output the results of the two. At the same time, the learning and updating module is used to correct the tracking and detection modules, which not only improves the classification and generalization level of the detection module, but also plays a great role in the stability of the algorithm. In short, the present invention is based on the tracking-detection-learning framework, optimizes and improves the algorithms of each module, solves the four main problems encountered in the long-term tracking problem, and has the advantages of high real-time performance, high robustness, and high detection accuracy.
Claims
1. An anti-occlusion mobile target tracking device based on ROI prediction and multi-module learning, It is characterized in that It includes tracking module, detection module, learning module and comprehensive module, among which: The tracking module includes a multi-feature extraction module and a correlation filtering module. A scale filter is added on the basis of a position filter as the algorithm framework of the tracking module. By adding a multi-feature extraction fusion optimization framework, PCA dimension reduction and QR decomposition are performed on the position filter and the scale filter. The detection module includes a ROI prediction module and a cascade classification module. The ROI prediction module uses a square root volume Kalman filter to estimate the position, and uses the estimated position to obtain the ROI area and use it as an input to the cascade classification module. The fHOG-SVM classifier is used in the cascade classification module; The tracking module and the detection module work synchronously, and the learning module calibrates, learns, and updates parameters with each other. The comprehensive module obtains the final output position through the coordination between multiple modules to achieve tracking of a single moving target. The multi-feature extraction module is obtained by splicing and fusing the M×N×1 grayscale feature and the M×N×31 fast directional gradient histogram, i.e., fHOG feature, wherein M and N are both positive integers, and the 31-dimensional fHOG feature is obtained by normalizing and truncating to obtain a 36-dimensional feature vector corresponding to each cell and then reducing the dimension using principal component analysis PCA, including 18-dimensional signed fHOG gradient, 9-dimensional unsigned fHOG gradient, and 4-dimensional features of normalized operations from the current cell and 4 domain cells in the diagonal domain; The basic framework of the correlation filter module is a discriminative scale space tracking algorithm framework, which uses position filters and scale filters to perform target positioning and scale evaluation in turn, and performs principal component analysis PCA dimensionality reduction and orthogonal triangular QR decomposition on the position filters and scale filters respectively to optimize the algorithm framework; The ROI prediction module estimates the ROI region by adding a square root volume Kalman filter; taking the target position predicted by the current frame as the center, defining a region with the aspect ratio of the previous frame and four times the area, and sending this region as the ROI region of the current frame to the cascade classification module; The cascade classification module includes an image element variance classifier, an fHOG-SVM classifier, and a nearest neighbor classifier, which are as follows: The ROI prediction module predicts the area where the target is most likely to appear in the current frame, and uses this as the input of the cascade classification module, namely the area to be detected. First, multiple sliding windows of different scales are generated in the area to be detected, and sent to the image element variance classifier to calculate the pixel grayscale variance of the window to be tested and the target frame image. The test sample with a variance less than half of the target sample variance is considered to be a negative sample; then the positive sample is used as the input of the fHOG-SVM classifier, the fHOG feature is extracted, and sent to the SVM classifier to obtain the positive and negative sample category results; finally, the positive sample window obtained by the first two classifiers is used as the input of the nearest neighbor classifier, and the similarity of each window with the online model is matched in turn, and the positive sample space of the online model is updated, so as to obtain the final positive sample of the cascade classification module, namely the output of the detection module.
2. A method for tracking moving targets with occlusion resistance based on ROI prediction and multi-module learning. It is characterized in that The following steps are involved: Step 1, multi-feature extraction: convert the input RGB three-channel image into a single-channel image, use the gamma correction method to standardize the image color space, calculate the image gradient, including the gradient value and gradient direction of each pixel, and then construct a 9-dimensional HOG feature vector; obtain the 36-dimensional feature vector corresponding to each cell through normalization truncation, extract 31-dimensional features through PCA dimensionality reduction, combine the features of each cell, obtain M×N×31-dimensional fHOG features from an M×N image, and splice them with the M×N×1 grayscale features to obtain M×N×32-dimensional fusion features; Step 2, correlation filter tracking: set the position filter and scale filter. For the position filter, first initialize the expected two-dimensional Gaussian output of the target, collect a sample with the target position as the center, use PCA dimension reduction to reduce the fusion feature from 32 dimensions to 18 dimensions, extract the 18-dimensional feature for each pixel of the sample, and multiply it by the two-dimensional Hamming window as the test input, and then use the inverse Fourier transform to determine the new position of the target; for the scale filter, first initialize the expected one-dimensional Gaussian output of the scale filter, extract samples at different scales with the target position as the center, pass each sample through the one-dimensional Hamming window as the test input, and then use the inverse Fourier transform to determine the new scale of the target; Step 3, ROI prediction area: Take the position of the target in the image at the previous moment as the observation value, use the square root volume Kalman filter algorithm to estimate the position of the target in the image at the next moment, delineate the area with the aspect ratio of the previous frame and four times the area, and send this area as the ROI area of the current frame to the detection module; Step 4, cascade classification: set the image element variance classifier, fHOG-SVM classifier, and nearest neighbor classifier. The ROI area is the area where the target of the current frame is most likely to appear. This ROI area is used as the input of the cascade classification module, that is, the area to be detected. First, multiple sliding windows of different scales to be tested are generated in the area to be tested, and sent to the image element variance classifier. The pixel grayscale variance of the window to be tested and the target frame image is calculated. The test sample with a variance less than half of the target sample variance is considered to be a negative sample; then the positive sample is used as the input of the fHOG-SVM classifier, the fHOG feature is extracted, and it is sent to the SVM classifier to obtain the positive and negative sample category results; finally, the positive sample window obtained by the first two classifiers is used as the input of the nearest neighbor classifier, and the similarity of each window with the online model is matched in turn, and the positive sample space of the online model is updated, so as to obtain the final positive sample of the cascade classification module, that is, the output of the detection module; Step 5, learning update: Use PN learning to optimize the performance of the classifier in the detection module in an online learning manner. In PN learning, first, the tracking module is used to predict the target position of the current frame. If the predicted position is detected as a negative sample by the detection module, the P expert will correct the positive sample that is incorrectly classified as a negative sample to a positive sample and send it to the training set. Then, the N expert compares the positive sample generated by the detection module with the positive sample obtained by the P expert, and selects the most credible sample as the output position. Step 6: Multi-module integration: Through the coordination between multiple modules, the final output position is obtained to achieve tracking of a single moving target.
3. The anti-occlusion moving target tracking method based on ROI prediction and multi-module learning according to claim 2, It is characterized in that In step 2, for the position filter, first initialize the expected two-dimensional Gaussian output of the target, collect a sample with the target position as the center, use PCA dimension reduction to reduce the fusion feature from 32 dimensions to 18 dimensions, extract the 18-dimensional feature for each pixel of the sample, and multiply it by the two-dimensional Hamming window as the test input, and then use the inverse Fourier transform to determine the new position of the target, as follows: First, assume that the target sample selected in the initial image is the positive sample f, and select a two-dimensional Gaussian function as the expected output sample g, so that the following formula is minimized: Among them, * represents the convolution operation, λ represents the weight coefficient, and f l represents the feature of the l-th channel, and h l represents the filter of the l-th channel, where l ∈ {1, 2,..., d} and d is the dimension of the selected features; Convert the above formula into complex frequency domain and solve it using Parseval formula: Among them, H l , F l , G are the corresponding variables obtained by performing the discrete Fourier transform (DFT) on h l , f l , g, and is the conjugate transpose of G; Use a training sample f t Update the filter parameters: Among them, and B t are the filter corresponding to a training sample f t for the numerator and denominator, and B t-1 are the filter numerator and denominator of the previous frame of this training sample, and η is the learning rate; If z t is an image sample, is a variable obtained by discrete Fourier transform, then the output y t is: Among them, and are the numerator and denominator of the filter in the previous frame; y t is the correlation score. By finding the maximum correlation score, the state estimation of the current target position is obtained.
4. The anti-occlusion moving target tracking method based on ROI prediction and multi-module learning according to claim 2, It is characterized in that In step 3, the position of the target in the image at the previous moment is taken as the observation value, and the square root volumetric Kalman filter algorithm is used to estimate the position of the target in the image at the next moment. The area is demarcated by the aspect ratio of the previous frame and four times the area, and this area is sent to the detection module as the current frame ROI area, as follows: For discrete nonlinear dynamic target tracking systems with additive noise: where x t and y t represent the state and measurement of the system at time t, respectively, f(·) and h(·) are the nonlinear state transition function and the nonlinear measurement function, respectively, the process noise w t-1 and the measurement noise v t-1 are independent of each other, and w t-1 ~N(0, Q t ), v t-1 ~N(0, R t ); State estimation includes time update and measurement update. When a fault is detected, the state parameters x t-1 and S t-1 of the last successful frame are used to initialize the filter, and then the filter gain K is calculated by the following formula t , the new state estimate and the square root factor S of the error covariance t : Among them, P xy,t is the cross-covariance matrix of the measurement prediction values, S yy,t is the square root of the self-covariance matrix of the measurement prediction values, is the system state prediction value, is the estimated measurement state prediction value, χ t , γ t are the weight matrices, S R,t is the square root of the covariance matrix of the measurement noise; Take the position v=(i,j) of the target in the image at the previous moment as the observation value, and estimate the position of the target in the image at the next moment. Define a region based on the aspect ratio of the previous frame and four times the area, and send this region as the current frame ROI region into the detection module.
5. The anti-occlusion moving target tracking method based on ROI prediction and multi-module learning according to claim 2, It is characterized in that In step 4, the SVM uses a kernel function to solve the non-linear problem. By establishing a hyperplane in the feature space as the decision surface, the isolation margin between positive and negative samples is maximized to separate the positive and negative samples. Assume the hyperplane is: wx + b = 0 where w represents the normal vector, determining the direction of the hyperplane, and b represents the offset, determining the distance between the hyperplane and the origin; The training sample set train = {(x 1 , y 1 ), (x 2 , y 2 ),..., (x n , y n )}, where x ∈ R n , y i ∈ {+1, -1}, i represents the i-th sample, and n represents the sample size; the classification surface needs to satisfy y i (wx i + b) ≥ 1, for i = 1, 2,..., m. The optimal hyperplane problem is transformed into: Introduce the Lagrangian function: should satisfy and Solve for the optimal solution The optimal weight method vector w * , the optimal offset b * : Therefore, the optimal hyperplane w * x + b * = 0, and the optimal classification function f(x) = sgn{w * x + b *}; Compare the similarity between the positive samples output by the fHOG-SVM classifier and the online model to achieve sample classification and update the positive sample space of the online model. The similarity calculation is as follows: Among them, S r is the relevant similarity, S + is the positive similarity, S - is the negative similarity, and is defined as follows: Among them, M represents the target model of the sample library, represents the positive sample, represents the negative sample, and p represents the sample to be measured; The formula for S is as follows: S(p i ,p j ) = 0.5(NCC(p i ,p j ) + 1) where NCC is defined as follows: Among them, μ i , σ i are the mean and standard deviation of the image block p i , μ j , σ j are the mean and standard deviation of the image block p j ; Finally, the calculated S r For comparison, S r The larger the value, the greater the possibility that the sample is the target. Set the threshold γ, S r Samples with a value greater than γ are considered positive samples, otherwise they are negative samples and are discarded. At the same time, new positive samples are added to the positive sample library of the online model for subsequent matching. The number of positive sample libraries of the online model is fixed. If the number is insufficient, additional samples will be added. If the number exceeds the upper limit, some samples will be randomly deleted and new samples will be added.
6. The anti-occlusion moving target tracking method based on ROI prediction and multi-module learning according to claim 2, characterized in that Step 6 is specifically as follows: The comprehensive module obtains the final output position through the coordinated work among multiple modules. According to the operation results of the detection module and the tracking module, it is divided into four cooperation modes: (1) Tracking success and detection success Detection success means that at least one sliding window passes through the detection module. After clustering the passed sliding windows, the final clustering result has only one clustering center. Detection failure means that no sliding window passes through the detection module, or multiple sliding windows pass through the detection module but the clustering result has multiple clustering centers; Tracking success means that the tracking module outputs a feature rectangle frame, and tracking failure means that no feature rectangle frame is output; If tracking is successful and detection is successful, cluster the detection results to obtain the relevant output results, and judge the overlap rate and credibility between the clustering center and the tracking module. If the overlap degree between the two is lower than the threshold of 0.5 and the credibility of the detection module is high, use the detection module to correct the result of the tracking module. If the overlap degree is higher than the threshold of 0.5, use the weighted average of the detection module and the tracking module as the result as the final output; (2) Tracking success and detection failure If tracking is successful but detection fails, directly use the output of the tracking module as the final output of the current frame; (3) Tracking failure and detection success If tracking fails but detection is successful, cluster the output sample frame of the detection module. If the final clustering result has only one clustering center, use the result of this clustering as the final output and re-initialize the tracking module with the result of this clustering, that is, the re-detection process when the target disappears and then enters again; If there are multiple clustering centers, it means that although it has passed the detection module, there are multiple different positions, and it is considered a detection failure; (4) Tracking failure and detection failure If both the tracking module and the detection module fail, it is considered that this detection is invalid and discarded.