Target tracking method based on multi-scale joint feature adaptive fusion
By adopting the adaptive fusion method of multi-scale joint features and contour-final scale processing in the target tracking algorithm, the problem of single feature extraction, non-adaptive fusion and filtering templates in the prior art is solved, and a more efficient and robust target tracking effect is achieved.
Patent Information
- Application Number
- CN202411862089.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing target tracking algorithm has problems such as single feature extraction methods, inability to adapt to feature fusion, and easy to learn irrelevant information in filtering templates.
Adaptive fusion method based on multi-scale joint features is adopted, combining HOG features, color histogram features, Haar local features and LBP features, and the impact of boundary effects is reduced through adaptive response fusion and contour exhaustive scale processing.
The target tracking algorithm's ability to characterize target features is improved, the processing ability of multi-scale targets is enhanced, and the occurrence of tracking drift and model pollution is reduced.
Smart Images

Figure CN119992125A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target tracking, and in particular relates to a target tracking method based on multi-scale joint feature adaptive fusion. Background Art
[0002] In recent decades, the field of computer vision has developed rapidly, and computers have acquired various image processing capabilities, including target detection, target recognition, target tracking, image segmentation, image super-resolution, and image fusion. Among them, target tracking has been widely used in many fields such as video surveillance and security, autonomous driving technology, smart home and Internet of Things, and medical image processing, and plays an important role in various fields. In-depth research on the principles and methods of target tracking technology is of great significance to promoting technological innovation and application in related fields.
[0003] In the practical application of existing target tracking algorithms, targets often have various attributes such as blur, deformation, rotation and low resolution.
[0004] There are three main problems with traditional tracking algorithms: First, the feature extraction method is single, among which the grayscale feature extraction method is simple but contains less information, the color feature is only effective when there is a large color difference between the target and the background, and the HOG feature has poor tracking robustness for blurred images; second, the feature fusion adopts a fixed ratio method and cannot adaptively adjust the feature weights according to the scene; third, due to the existence of boundary effects, the filter template is easy to learn irrelevant information. Summary of the invention
[0005] In view of this, the main purpose of the present invention is to provide a target tracking method based on multi-scale joint feature adaptive fusion.
[0006] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0007] An embodiment of the present invention provides a target tracking method based on multi-scale joint feature adaptive fusion, the method comprising:
[0008] Step 1, input the first frame image, and initialize the feature filter, scale filter and color histogram model;
[0009] Step 2, input the t-th frame image, perform contour scale detection, extract and screen candidate frames, perform extreme scale scaling on the screened candidate frames to obtain the final candidate frames, and extract the candidate target area from the final candidate frames;
[0010] Step 3, extract features for each candidate target region, and determine the HOG features, color histogram features, Haar local features, and LBP feature responses of each candidate target region;
[0011] Step 4, calculating the adaptive fusion parameters, performing adaptive response fusion, and obtaining the final response result; determining the final target prediction candidate area according to the maximum value of the final response result, returning the candidate target area corresponding to the maximum response, and determining the scale of the target area as the target scale of the current frame;
[0012] Step 5, using the final target prediction candidate region of the t-th frame image to construct a basic training sample, training and updating the filter model and the color histogram model, and obtaining the filter model and the color histogram model of the next frame image;
[0013] Step 6: Repeat steps 2 to 5 until the tracking is completed.
[0014] In the above scheme, the first frame image is input to initialize the feature filter, the scale filter and the color histogram model, specifically including:
[0015] The target area of the first frame image has been manually marked. The target area is used as the basic sample and a training sample is constructed by cyclic shifting.
[0016] Using the ridge regression method, the initialization filter model ω is obtained through the training samples. Where a represents the basis vector corresponding to the sample, b represents the standard response corresponding to the sample, ω represents the filter model parameter to be solved, and λ represents the regularization coefficient;
[0017] The kernel function method is used to optimize the filter model ω, and the optimized solution of the filter model is obtained as In the formula, Represents the calculation of the correlation between the sample and its own kernel, and I represents the unit matrix;
[0018] Obtaining an initialized color histogram model using the target area as a basic sample;
[0019] An initialized scale filter model is obtained using the target area as a basic sample.
[0020] In the above scheme, the input t-th frame image is subjected to contour scale detection, candidate frames are extracted and screened, the screened candidate frames are subjected to extreme scale scaling to obtain a final candidate frame, and a candidate target area is extracted from the final candidate frame, specifically including:
[0021] The contour of the target area of the t-1th frame image is extracted, and the edge calculation is performed on each pixel point to obtain the edge response value of each point, and then the contour sparse response map is obtained by the non-maximum suppression algorithm; each pixel point a in the contour sparse response map has a contour value size l a and contour direction β a , for each point's contour value la Make a judgment, a If it is greater than 0.15, the point is considered as a contour and is retained;
[0022] Contour clustering is performed on the contour sparse response map, and the contour responses of the 8-connected domains around each pixel in the target area are compared. If the contour directions β of two pixels are a If the difference is less than 90°, they are judged to be the same type of contour, otherwise they are not the same type of contour;
[0023] The correlation of the contour clustering results is judged. For the center coordinates (x i ,y i ) of the contour class e i With the center coordinates (x j ,y j ) of the contour class e j ,pass Determine the correlation between two contour classes, E c (e i ,e j ) represents the correlation between two contour classes, β ij is the coordinate (x i ,y i ) and coordinates (x j ,y j ), α c Represents the relevant parameters, Δ d Represents the shortest distance between two classes;
[0024] Determine each contour class e i The weight value ξ b , S eb Represents the candidate bounding box b wh With contour class i The area of the region included, S e Represents the contour class e i Area covered;
[0025] According to each contour class e i The weight value ξ b Determine each contour class e i The final confidence p of the candidate bounding box b b ;
[0026] Extract candidate bounding boxes and retain the top 35% of the candidate bounding boxes with confidence;
[0027] Compare the area of the candidate bounding boxes with the top 35% confidence level with the target area determined by the t-1 frame image, and retain the candidate bounding boxes that are larger than 0.6 times and smaller than 1.4 times the target area of the t-1 frame image;
[0028] The candidate target areas are extracted by scaling the candidate regions retained by the extreme scale. The center point of each candidate region calibrated by the candidate box is (x, y). The target width and height determined by the t-1 frame image are w and h respectively. Each candidate region is scaled to a certain scale to extract candidate target image regions of i scales.
[0029] In the above scheme, the input t-th frame image first extracts the gradient histogram (HOG) features of the candidate target area, including the following specific steps:
[0030] The RGB image of the candidate area is converted into a grayscale image by the following formula;
[0031] Perform Gamma correction on the grayscale converted image;
[0032] By using [-1,0,1] as the horizontal gradient operator, [-1,0,1] T As a vertical gradient operator, determine the horizontal image gradient vector and the vertical image gradient vector of the corrected image;
[0033] Determine the gradient amplitude |G(x,y)| and the gradient direction θ(x,y) according to the horizontal image gradient vector and the vertical gradient vector;
[0034] The corrected image is divided into small cell units. The gradient direction in each cell is divided into 9 parts from 0° to 180°. Every 20° direction is represented by a direction area. 1 to Z 9 Represent 9 directional areas respectively, perform weighted statistics on the gradients of the 9 parts, and obtain the gradient direction histogram of the cell, that is, the feature vector of 9 length;
[0035] Several adjacent cells are synthesized and merged into a block. In each block, the feature vectors of all cells in it are connected in series to form the HOG feature of the current block. The gradient histogram in the block is normalized to obtain the normalized gradient amplitude.
[0036] The HOG features in all blocks are integrated to obtain 36-dimensional HOG features.
[0037] In the above solution, determining the color histogram features of each candidate target area specifically includes:
[0038] According to the Bayesian statistical formula, the likelihood probability that pixel x belongs to the target is
[0039]
[0040] Where O is the foreground area of the candidate target area, and its surrounding space is the background area S. I is the target candidate area composed of the foreground area O and the background area S. b represents the classification interval of the color histogram. x Represents the combination of RGB values of pixel x in interval b;
[0041] The likelihood term is estimated for the likelihood probability. In the formula, Indicates the number of pixels in the b classification interval in the O region in I, represents the number of pixels in the b classification interval in the S region of I, |O| represents the area of the foreground area, and |S| represents the area of the background area;
[0042] Convert the likelihood probability of pixel point x belonging to the target into Here, 0.5 represents the estimated response value.
[0043] In the above solution, determining the Haar local features of each candidate target area specifically includes:
[0044] The Haar local features include edge Haar features, linear Haar features, center Haar features and diagonal Haar features;
[0045] Constructing the integral graph Among them, the pixel value ii(x,y) at the position (x,y) of the candidate area is the sum of the grayscale values i(x',y') of all pixels in the upper left corner of the original image (x,y), s(x,y) represents the cumulative sum in the row direction, s(x,-1)=0 is initialized, ii(x,y) represents an integral image, ii(-1,y)=0 is initialized;
[0046] Scan the image line by line, recursively calculate the cumulative sum s(x,y) of each pixel (x,y) in the row direction and the value of the integral image ii(x,y) are s(x,y)=s(x,y-1)+i(x,y) respectively 、 ii(x,y)=ii(x-1,y)+s(x,y) ;
[0047] Traverse the candidate target area, and when the pixel in the lower right corner of the image is scanned, the image integral map ii (x, y) is constructed;
[0048] The Haar eigenvalues in the candidate target area are calculated based on the integral image.
[0049] The Haar features include 2 edge Haar features, 1 center Haar feature, 4 linear Haar features and 1 diagonal Haar feature;
[0050] To calculate the edge Haar feature, it is necessary to search the integral image 6 times, to calculate the line feature, it is necessary to search 8 times, to calculate the center Haar feature, it is necessary to search 8 times, and to calculate the diagonal Haar feature, it is necessary to search 9 times. The feature extraction is performed in the form of a sliding window, sliding 1 pixel unit each time. For each sliding window, the 8 Haar features in the window are calculated.
[0051] In the above solution, determining the LBP features of each candidate target area specifically includes:
[0052] Select a 3×3 neighborhood of each pixel in the candidate target area, take the gray value of the central pixel as the reference, and compare the gray values of the pixels in the surrounding 8 neighborhoods. If the value of the neighborhood pixel is greater than or equal to the value of the central pixel, mark the position as 1, otherwise it is marked as 0. The LBP feature calculation formula is expressed as In the formula, (x c ,y c ) represents the center pixel coordinates, p represents the pth pixel in the neighborhood, i c Represents the gray value of the center pixel, i p represents the gray value of the neighborhood pixel, s(·) is the sign function, expressed as
[0053] In the above scheme, the calculation of the adaptive fusion parameters, the adaptive response fusion, and the final response result are obtained; and the final target prediction candidate area is determined according to the maximum value of the final response result, which specifically includes:
[0054] The four feature response weights were evaluated by the average peak correlation energy (APCE). In the formula, F max Indicates the maximum value in the response graph, F min represents the minimum value in the response graph, (w,h) represents the response coordinates; Step 4 obtains the response of each feature, which is recorded as r es1 、r es2 、r es3 、r es4 , calculate the APCE of each characteristic response as P APCE1 , P APCE2 , P APCE3 , P APCE4 ;
[0055] Construct a mapping function to make the APCE ratio of each response float within a reasonable range. The mapping function is: In the formula, ω ss represents the pre-fusion parameter of the response, γ m As a hyperparameter, the parameter ω can be ss Keep it within a reasonable range;
[0056] Determine the parameter ω of each target area image ss And normalize it, the obtained fusion parameter ω s for
[0057] According to the fusion parameter ω s Determine the final fusion response of the four features of each target area as r es =ω s1 × es1 +ω s2 × es2 +ω s3 × es3 +ω s4 × es4 ; In the formula, r es Represents the final response result, returns the candidate target area corresponding to the maximum response result as the final target candidate area, and sets the scale of the target area (w i ,h i ) is determined as the target scale for the current frame.
[0058] In the above scheme, the basic training samples are constructed, and the filter model and the color histogram model are trained and updated to obtain the filter model and the color histogram model of the next frame image; specifically, the following steps are included:
[0059] The final target candidate area is used as the basic sample to construct the training sample.
[0060] The training samples and the target candidate regions form a kernel matrix, K z =C(k az );where K z represents the kernel matrix composed of training samples and target candidate regions, k az Represents the kernel correlation calculation between the training sample a and the sample z in the area to be detected;
[0061] Determine the objective function expression according to the kernel matrix Get the predicted response R(z) in time domain, R(z) = K z α; where K z represents the kernel matrix composed of training samples and target candidate regions, α represents the correlation filter model coefficient of the previous frame image, and the target position is the position corresponding to the maximum value of R(z);
[0062] Update the color histogram model x and the coefficient α of the filter model, x 1:t =(1-η)x 1:t-1 +ηx t , α 1:t =(1-η)α 1:t-1 +ηα t; In the formula, η represents the learning rate, x t represents the target sample vector of the current frame, x 1 :t represents the target sample vector learned from the beginning to time t, α t Represents the filter coefficient of the current frame, α 1 :t represents the filter coefficients learned from the beginning to time t.
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] (1) To address the problems of single feature extraction method and inability to highlight important features during feature fusion, a joint feature extraction method is proposed. By introducing the Haar feature in the field of face detection, a Haar local feature is proposed, and an adaptive feature fusion method is designed to adaptively fuse HOG features, color histogram features, Haar local features and LBP features to highlight important features and suppress features with low discrimination, further improving the algorithm's ability to represent target features.
[0065] (2) In order to solve the problem that the filter template is easy to learn irrelevant background information, a spatial restriction model is introduced to reduce the impact of boundary effects. Due to the introduction of the spatial restriction model, the ridge regression least squares method cannot meet the optimization requirements of the algorithm, so the alternating direction multiplier method is used for solution optimization.
[0066] (3) To address the problem of scale changes during tracking, the contour extreme scale processing method is used, and the contour scale detection method is introduced into the tracking field. It is combined with the extreme scale processing method to improve the algorithm's ability to handle multi-scale targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings described herein are used to disclose a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0068] Figure 1 It is the Bird1 sequence tracking result;
[0069] Figure 2 It is the BlurOwl sequence tracking result;
[0070] Figure 3 It is the Box sequence tracking result;
[0071] Figure 4 It is the Panda sequence tracking result;
[0072] Figure 5 It is the DragonBaby sequence tracking result;
[0073] Figure 6 It is the Human6 sequence tracking result;
[0074] Figure 7 It is the Human7 sequence tracking result;
[0075] Figure 8 It is the Ironman sequence tracking result;
[0076] Fig. 9 It is the accuracy and success rate curve after summarizing the data of each algorithm;
[0077] Fig.10 This is a schematic diagram of 8 types of Haar feature blocks. DETAILED DESCRIPTION
[0078] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0079] The present invention provides a target tracking method based on multi-scale joint feature adaptive fusion, and the implementation steps are as follows:
[0080] Step 1: Input the first frame image of the video sequence and initialize the filter model and color histogram model.
[0081] Step 2: Input the tth (t>1) frame image, first perform contour scale detection, extract and screen candidate borders, then perform extreme scale scaling to obtain the final candidate borders and extract the candidate target area.
[0082] Step 3, extracting features from each candidate target region obtained in step 2, and calculating the HOG features, color histogram features, Haar local features, and LBP feature responses of each candidate target region;
[0083] Step 4, calculate the adaptive fusion parameters, perform adaptive response fusion on the four feature responses obtained in step 4, and obtain the final response result; determine the final target prediction candidate area according to the maximum value of the response result; return the candidate target area corresponding to the maximum response, and determine the scale of the target area as the target scale of the current frame;
[0084] Step 5, using the final target prediction candidate region of the t-th frame image to construct a basic training sample, training and updating the filter model and the color histogram model, and obtaining the filter model and the color histogram model of the next frame image;
[0085] Step 6: Repeat steps (2) to (5) until the tracking is completed.
[0086] In step 1, the first frame image of the video sequence is input, and the filter model and the color histogram model are initialized, including the following specific steps:
[0087] Step 1.1: The target area of the first frame image has been manually marked before being input into this algorithm. The target area is used as the basic sample and a training sample is constructed by cyclic shift. The specific operation method is as follows:
[0088] (1) Assume that the size of the input image I is i×j, expressed as
[0089]
[0090] (2) Construct permutation matrices P and Q to enable the image to be quickly translated. The matrices P and Q are in the form of
[0091]
[0092] The sizes of P and Q are both i×j
[0093] (3) Matrix P is responsible for up and down translation. Matrix P is multiplied by I once on the left, and I is shifted down one row. Matrix Q is responsible for left and right translation. Matrix Q is multiplied by I once on the right, and I is shifted right one column. The formula is expressed as
[0094] {P l IQ r |l∈[1,i],r∈[1,j]}
[0095] In the formula, l and r represent the number of up and down translation and left and right translation. The basic sample is constructed into a circulant matrix by the above cyclic shift operation as a training sample.
[0096] Step 1.2, using the ridge regression method, initialize the filter model ω through the training samples obtained in step 1.1. The specific operation method is as follows:
[0097] (1) Let the training sample matrix obtained in step 1.1 be A, x be the unknown vector, and Ax be the model prediction value; let b be the standard response corresponding to the sample, that is, the real observation data. The ideal relationship between the training sample and the standard response is
[0098] Ax=b
[0099] (2) In the actual tracking process, there is an error between the model prediction value and the observed data. In order to minimize the sum of squared errors, a correction vector Δb is introduced, and the solution of the original equation is converted to the minimum value of the correction vector Δb. At this time, the solution to the equation Ax = b is equivalent to
[0100]
[0101] (3) Adding a regular term to the above equation makes it have a unique analytical solution and prevents overfitting. The objective function of ridge regression is:
[0102]
[0103] In the formula, f(·) represents the regression function, a represents the basis vector corresponding to the sample, b represents the standard response corresponding to the sample, ω represents the filter model parameter to be solved, and λ represents the regularization coefficient. Taking the minimum value and expanding it to the complex domain, we can get
[0104]
[0105] Where H represents the conjugate transpose.
[0106] Use Fourier diagonalization to H A is calculated and we get Substituting into the above formula, the filter model parameter ω is
[0107]
[0108] Perform Fourier transform on both sides of the above equation and simplify it using the diagonal matrix property to obtain
[0109]
[0110] Step 1.3, use the kernel function method to optimize the filter model. The specific operation method is as follows:
[0111] (1) The parameter to be solved ω is transformed into the dual space α (α = {α 1 ,α 2 ,...,α N}),have to
[0112]
[0113] At this time, the result variable of the solution changes from ω to α, α N represents the mapping weight of the Nth sample in the low-dimensional space in the high-dimensional space, a N Represents the basis vector corresponding to the sample
[0114] (2) The objective function can be rewritten from the above formula as
[0115]
[0116] In the formula, K(z,a N ) represents the training samples in high-dimensional space and The regression function becomes
[0117]
[0118] The calculated filter model optimization solution is
[0119]
[0120] In the formula, Represents the calculation of the kernel correlation between the sample and itself, and I represents the identity matrix.
[0121] Step 1.4, using the target area as the basic sample, initialize the color histogram model. The specific operation method is as follows:
[0122] (1) The foreground area of the target area is O, and the surrounding space is the background area S. The foreground area O and the background area S together constitute the target candidate area I. According to the Bayesian statistical formula, the likelihood probability of pixel point x belonging to the target is
[0123]
[0124] In the formula, b represents the classification interval of the color histogram, b x Represents the combination of RGB values of pixel x in interval b.
[0125] (2) Estimation of likelihood term for likelihood probability is performed using the following method:
[0126]
[0127] In the formula, Indicates the number of pixels in the b classification interval in the O region in I, represents the number of pixels in the b classification interval in the S region of I, |O| represents the area of the foreground area, and |S| represents the area of the background area
[0128] (3) According to the likelihood estimation, the likelihood probability that the pixel point x in the target candidate area belongs to the target can be expressed as
[0129]
[0130] Where 0.5 represents the estimated response value. If the RGB combination of pixel x is x , which does not appear in O and S, has a probability of 0.5.
[0131] In step 2, input the t(t>1)th frame image, first perform contour scale detection, extract and screen candidate borders, then perform extreme scale scaling to obtain the final candidate borders, and finally extract the candidate target area. The specific operation method is as follows:
[0132] Step 2.1, perform contour extraction based on the target area of the previous frame. Perform edge calculation on each pixel point of the input image to obtain the edge response value of each point, and then obtain the contour sparse response map through the non-maximum suppression algorithm. Each pixel point a in the contour sparse response map has a contour value size l a and contour direction β a , for each point's contour value l a Make a judgment, a If it is greater than 0.15, the point is judged as a contour and retained.
[0133] Step 2.2: Perform contour clustering on the contour response obtained in step 2.1, and compare the contour responses of the 8-connected domains around each pixel in the target area. If the contour directions β of two pixels are a If the difference is less than 90°, they are judged to be the same type of contour, otherwise they are not the same type of contour.
[0134] Step 2.3, perform correlation judgment on the contour class obtained in step 2.2. i ,y i ) of the contour class e i With the center coordinates (x j ,y j ) of the contour class e j , the correlation between two contour classes is calculated using the following formula
[0135]
[0136] E c (e i ,e j ) represents the correlation between two contour classes, β ij is the coordinate (x i ,y i ) and coordinates (x j ,y j ), α c Represents the relevant parameters, Δ d Represents the shortest distance between two classes.
[0137] Step 2.4, calculate the contour class weight. Calculate each contour class e according to the following formula i The weight value ξ b
[0138]
[0139] S eb Represents the candidate bounding box b wh With contour class i The area of the region included, S e Represents the contour class ei Area included
[0140] Step 2.5, calculate the confidence of the candidate border. The specific operation method is as follows:
[0141] (1) Calculate each contour class e obtained in step 2.2 according to the following formula i The sum of the contour values of all pixel points a in m i
[0142]
[0143] (2) Calculate each contour class e according to the following formula i The confidence p of the candidate bounding box b bb
[0144]
[0145] sb w With b h Respectively represent the width and height of the candidate border b, σ b Set the value to 1.5.
[0146] (3) Suppress the central contour class and highlight the outer contour class. The contour class e is obtained according to the following formula i The final confidence of the candidate bounding box is
[0147]
[0148] Where b in The width and height are b w / 2 and b h / 2 .
[0149] Step 2.6, extract candidate bounding boxes and retain the bounding boxes with the top 35% confidence levels.
[0150] Step 2.7, candidate border screening. Calculate the area of the candidate border region and compare it with the target region area determined in the previous frame. Retain the candidate borders that are larger than 0.6 times and smaller than 1.4 times the target region area of the previous frame.
[0151] Step 2.8, use the extreme scale scaling to extract the candidate target area. The center point of each candidate area calibrated in step 2.7 is (x, y), and the width and height of the target determined in the previous frame are w and h respectively. Each candidate area is scaled to a certain scale to extract the candidate target image area of i scales. The center of these images is still (x, y), and the width and height are w and h respectively. 1 ~w i and h 1 ~h i .
[0152] In step 3, features are extracted for each candidate target region obtained in step 2, and the HOG features, color histogram features, Haar local features, and LBP feature responses of each candidate target region are calculated. The specific steps are as follows:
[0153] Step 3.1, extract the HOG features of the candidate target area. The specific steps are as follows:
[0154] (1) Image grayscale
[0155] The RGB image of the candidate area is converted into a grayscale image by the following formula
[0156] I gray =0.3×R c +0.59×G c +0.11×B c
[0157] In the formula, I gray Represents the converted grayscale image, R c , G c , B c Represent the red, green and blue components of the original image respectively.
[0158] (2) Gamma Correction
[0159] Gamma correction is performed on the grayscale converted image to normalize the pixel values of the image. The correction formula is:
[0160]
[0161] In the formula, γ g represents the compression factor, which is generally set to 0.5. (x, y) represents the pixel point. H g (x,y) represents the correction result image
[0162] (3) Calculate image gradient
[0163] Use [-1,0,1] as the horizontal gradient operator, [-1,0,1] T As a vertical gradient operator, the image gradient is calculated. The horizontal gradient of the image is calculated by the following formula
[0164] G x (x,y)=H g(x+1,y) -H g(x-1,y)
[0165] In the formula, G x (x, y) represents the horizontal gradient vector. The vertical gradient of the image is calculated by the following formula
[0166] G y(x,y)=H(x,y+1)-H(x,y-1)
[0167] In the formula, G y (x,y) represents the vertical gradient vector.
[0168] The gradient magnitude |G(x,y)| and the gradient direction θ(x,y) are calculated by the following two equations respectively:
[0169]
[0170] (4) Constructing gradient direction histogram
[0171] The corrected image is divided into small cell units (also called Cell), and the gradient direction in each cell is divided into 9 parts from 0° to 180°. Every 20° direction is represented by a direction area. 1 to Z 9 Represent 9 directional regions respectively. Perform weighted statistics on the gradients of the 9 parts to obtain the gradient direction histogram of the cell, that is, the 9-length feature vector.
[0172] (5) Synthesize Cell into Block
[0173] Several adjacent cells are combined into a block. In each block, the feature vectors of all cells in it are concatenated to form the HOG feature of the current block. The gradient histogram in the block is normalized. The normalized gradient amplitude is
[0174]
[0175] In the formula, g i represents the initial value of the gradient, and ε represents a very small constant to prevent the denominator from being zero.
[0176] (6) Integrate HOG features
[0177] The HOG features in all blocks are integrated to obtain multi-dimensional HOG features.
[0178] Step 3.2, calculate the color histogram features of the candidate target area. The specific operation method is as follows:
[0179] (1) Let the foreground region O of the candidate target region and its surrounding space be the background region S. The foreground region O and the background region S together constitute the target candidate region I. According to the Bayesian statistical formula, the likelihood probability of pixel point x belonging to the target is
[0180]
[0181] In the formula, b represents the classification interval of the color histogram, bx Represents the combination of RGB values of pixel x in interval b.
[0182] (2) Estimation of likelihood term for likelihood probability is performed using the following method:
[0183]
[0184] In the formula, Indicates the number of pixels in the b classification interval in the O region in I, It represents the number of pixels in the b classification interval in the S region of I, |O| represents the area of the foreground area, and |S| represents the area of the background area.
[0185] (3) Convert the likelihood probability of pixel point x belonging to the target into
[0186]
[0187] Where 0.5 represents the estimated response value. If the RGB combination of pixel x is x , which does not appear in O and S, has a probability of 0.5
[0188] Step 3.3, using the integral image method to calculate the four Haar local features of the candidate target area, namely edge Haar features, linear Haar features, center Haar features and diagonal Haar features. The specific steps are as follows:
[0189] (1) Construct an integral graph. The definition formula of the elements in the integral graph is:
[0190]
[0191] The above formula indicates that the pixel value ii(x,y) at the candidate region position (x,y) is the sum of all pixel grayscale values i(x',y') in the upper left corner of the original image (x,y). s(x,y) represents the cumulative sum in the row direction, and s(x,-1) is initialized to 0. ii(x,y) represents an integral image, and ii(-1,y) is initialized to 0.
[0192] (2) Scan the image line by line, recursively calculate the cumulative sum s(x,y) of each pixel (x,y) in the row direction and the value of the integral image ii(x,y) respectively
[0193] s(x,y)=s(x,y-1)+i(x,y)
[0194] ii(x,y)=ii(x-1,y)+s(x,y)
[0195] Traversing the candidate target area, when scanning the pixel in the lower right corner of the image, the image integral map ii(x,y) is constructed.
[0196] (3) Calculate the Haar feature value in the candidate target area according to the integral image. Eight Haar features are selected, including two edge Haar features, one center Haar feature, four linear Haar features and one diagonal Haar feature.
[0197] The Haar feature is calculated by subtracting the sum of the pixel values in the black area from the sum of the pixel values in the white area. Since the integral image is the sum of the grayscale values of all pixels in the upper left corner of a pixel, the sum of the pixels in any rectangular area of the image can be calculated through the integral image, and then any Haar feature value can be calculated.
[0198] Calculating edge Haar features requires searching the integral graph 6 times, calculating line features requires searching 8 times, calculating center Haar features requires searching 8 times, and calculating diagonal Haar features requires searching 9 times. The feature extraction is performed in the form of a sliding window, sliding 1 pixel unit each time, and for each sliding window, the 8 types of Haar features in the window are calculated.
[0199] Step 3.4, calculate the LBP features of the candidate target area.
[0200] The LBP feature uses a vector histogram to represent the feature. The calculation method is as follows: select a 3×3 neighborhood of each pixel in the candidate target area, take the grayscale value of the central pixel as the reference, and compare the grayscale values of the pixels in the surrounding 8 neighborhoods. If the value of the neighborhood pixel is greater than or equal to the value of the central pixel, mark the position as 1, otherwise it is marked as 0. The LBP feature calculation formula is expressed as
[0201]
[0202] In the formula, (x c ,y c ) represents the center pixel coordinates, p represents the pth pixel in the neighborhood, i c Represents the gray value of the center pixel, i p represents the gray value of the neighborhood pixel, s(·) is the sign function, expressed as
[0203]
[0204] In step 4, the adaptive fusion parameters of the four features obtained in step 3 are calculated, and adaptive response fusion is performed to obtain the final response result of each target area. According to the maximum value of the response result, the final target prediction candidate area is determined. The specific operation method is as follows:
[0205] Step 4.1, use the average peak correlation energy (APCE) to evaluate the four characteristic response weights. The APCE calculation formula is:
[0206]
[0207] In the formula, F max Indicates the maximum value in the response graph, F min represents the minimum value in the response graph, and (w,h) represents the response coordinates. Step 4 obtains the response of each feature, which is recorded as r es1 、r es2 、r es3 、r es4 , calculate the APCE of each characteristic response as P APCE1 , P APCE2 , P APCE3 , P APCE4
[0208] Step 4.2, construct a mapping function to make the APCE ratio of each response float within a reasonable range. The mapping function is
[0209]
[0210] In the formula, ω ss represents the pre-fusion parameter of the response, γ m As a hyperparameter, the parameter ω can be ss Control within a reasonable range.
[0211] Step 4.3, calculate the parameter ω of each target area image ss The value of and normalize it to get the fusion parameter ω s The formula is
[0212]
[0213] Step 4.4, according to the fusion parameter ω obtained in step 4.3 s , the final fusion response of the four features in each target area is
[0214] r es =ω s1 × es1 +ω s2 × es2 +ω s3 × es3 +ω s4 × es4
[0215] In the formula, r es represents the final response result, returns the candidate target area corresponding to the maximum response result as the final target candidate area, and converts the scale (w i ,h i ) is determined as the target scale for the current frame.
[0216] In step 5, a basic training sample is constructed, and the filter model and the color histogram model are trained and updated to obtain the filter model and the color histogram model of the next frame image. The specific processing steps are as follows:
[0217] Step 5.1, construct training samples using the final target candidate region obtained in step 4.4 as the basic sample. The specific steps are as follows.
[0218] (1) Assume that the size of the input image I is i×j, expressed as
[0219]
[0220] (2) Construct permutation matrices P and Q to enable the image to be quickly translated. The matrices P and Q are in the form of
[0221]
[0222] The sizes of P and Q are both i×j
[0223] (3) Matrix P is responsible for up and down translation. Matrix P is multiplied by I once on the left, and I is shifted down one row. Matrix Q is responsible for left and right translation. Matrix Q is multiplied by I once on the right, and I is shifted right one column. The formula is expressed as
[0224] {P l IQ r |l∈[1,i],r∈[1,j]}
[0225] In the formula, l and r represent the number of up and down translation and left and right translation. The basic sample is constructed into a circulant matrix by the above cyclic shift operation as a training sample.
[0226] Step 5.2: The training samples obtained in step 5.1 and the target candidate regions form a kernel matrix. The calculation method is as follows:
[0227] K z =C(k az )
[0228] In the formula, K z represents the kernel matrix composed of training samples and target candidate regions, k az Represents the kernel correlation calculation between the training sample a and the sample z in the area to be detected.
[0229] Substitute the above formula into the objective function expression The predicted response R(z) in time domain can be obtained, which is expressed as
[0230] R(z)=K z α
[0231] In the formula, K zrepresents the kernel matrix composed of training samples and target candidate areas, α represents the correlation filter model coefficient of the previous frame image, and the target position is the position corresponding to the maximum value of R(z)
[0232] Step 5.3, update the color filter model x and the coefficient α of the related filter, specifically:
[0233] x 1:t =(1-η)x 1:t-1 +ηx t
[0234] α 1:t =(1-η)α 1:t-1 +ηα t
[0235] In the formula, η represents the learning rate, x t represents the target sample vector of the current frame, x 1 :t represents the target sample vector learned from the beginning to time t, α t Represents the filter coefficient of the current frame, α 1 :t represents the filter coefficients learned from the beginning to time t.
[0236] In step 6, steps (2) to (5) are repeated until the tracking is completed.
[0237] The effects of the present invention are further described below in conjunction with simulation experiments.
[0238] Figures 1 to 8 The experimental results of the algorithm proposed in this invention and the comparison algorithm in each sequence are shown, where the algorithm is represented by JFAFMT (Joint Feature Adaptive Fusion Multiscale Tracking). The selected comparison algorithms are all correlation filter tracking algorithms, including CSK, CN2, DSST, IBCCF and Staple.
[0239] The Bird1 sequence results are as follows Figure 1 As shown in the figure, the target in this sequence has the characteristics of deformation, rapid movement and occlusion. At the 15th frame, the CSK and CN2 algorithms have tracking drift, and learn with the bird's wings as the target. Then, the bird flaps its wings, causing the two algorithms to fail to track. At the 126th frame, heavy fog began to appear, and the target was largely obscured by the fog. The target became blurred and the features were not obvious. The DSST, IBCCF and Staple algorithms began to have slight tracking drift. By the 182nd frame, only the algorithm in this chapter tracked correctly, and the rest of the algorithms all lost the tracking target.
[0240] BlurOwl sequence results are as follows Figure 2As shown in the figure, the target in this sequence has the characteristics of in-plane rotation, motion blur, rapid motion and scale change. At the 48th frame, the target begins to move downward rapidly and becomes blurred. The CN2 algorithm cannot track the target correctly, and the detection frame does not move with the target. At the 383rd frame, the target begins to shake more violently and is very blurred. All algorithms have a certain degree of drift, but the drift of the algorithm in this chapter is smaller. By the 438th frame, the target's jitter has weakened, and only the algorithm in this chapter accurately tracks the target again. The other algorithms all stay in place and cannot track the target area again.
[0241] Box sequence results are as follows Figure 3 As shown in the figure, the target in this sequence has characteristics such as occlusion, out-of-plane rotation, background clutter and illumination changes. At the 119th frame, the IBCCF algorithm learned too many can features, the sample model was polluted, and the target was gradually lost. At the 225th frame, the target moved to the upper left corner area, and the book on the left had similar color features to the target. The CN2 algorithm lost the target and stayed on the left side of the book. At the 484th frame, the target passed behind the caliper and there was occlusion. The white and yellow text of the target was occluded, and the caliper color was black. The area below was also black due to dimming of the light. In addition to the algorithms in this chapter, the CSK, DSST and Staple algorithms all have a certain degree of drift.
[0242] Panda sequence results are as follows Figure 4 As shown in the figure, the targets in this sequence have characteristics such as low resolution, scale change, deformation, and in-plane rotation. In frame 143, the panda walks to the right and suddenly turns around. The DSST and IBCCF algorithms cannot follow and adjust in time. Only the panda's head is left in the detection frame, and the model is contaminated. In frame 413, the panda moves to the upper left corner and no longer faces the camera from the side. The CSK algorithm shifts to the whiteboard area on the left. Since the models of the DSST and IBCCF algorithms are contaminated, they no longer follow the target movement. In frame 594, the panda returns to the right area. At this time, the Staple algorithm can only track the head, and only the algorithm in this chapter still tracks all parts of the target.
[0243] The DragonBaby sequence results are as follows Figure 5 As shown in the figure, the target in this sequence has the characteristics of in-plane rotation, out-of-plane rotation, rapid movement and out of field of view. In the 28th frame, the child in yellow clothes turns around, and the CSK and DSST algorithms drift and lose the tracking target. In the 41st frame, the CN2 algorithm has drifted to the yellow clothes area. The child is hit hard by the monster dragon, and the target begins to move backwards quickly. The IBCCF algorithm also partially loses the target. In the 48th frame, the child raises his hands and kicks the monster, and part of the target area is blocked. At this time, only the algorithm in this chapter can continue to track the target.
[0244] The Human6 sequence results are as follows Figure 6 As shown in the figure, the targets in this sequence have characteristics such as scale change, occlusion, deformation and rapid motion. In the 232nd frame, the target is enlarged, and the CSK, CN2 and IBCCF algorithms cannot adapt to the scale change of the target and drift, losing the tracking target. In the 357th frame, the pedestrian is blocked by the road sign, the target disappears, and the DSST, Staple and this algorithm begin to learn the features of the road sign. By the 598th frame, the target moves to other areas, and the DSST algorithm stays on the road sign. Only the algorithm in this paper and the Staple algorithm can accurately track it.
[0245] The Human7 sequence results are as follows Figure 7 As shown in the figure, the target in this sequence has characteristics such as illumination change, scale change, motion blur and occlusion. In the 58th frame, all algorithms can track the target, but CSK, CN2 and IBCCF cannot fully adapt to the scale change of the target, and the model information is polluted. In the 121st frame, the target continues to move forward and the scale becomes smaller. At this time, the CN2, DSST and IBCCF algorithms all lose the target. In the 250th frame, the target moves to the left side of the stairs, and the CSK algorithm also loses the target. Only the algorithm in this paper and the Staple algorithm can accurately track the target.
[0246] The Ironman sequence results are as follows Figure 8 As shown in the figure, the targets in this sequence have characteristics such as illumination change, in-plane rotation, occlusion, and background clutter. In the 26th frame, the overall environment changes from bright to dark, and Iron Man rotates to the right, causing the Staple algorithm to drift and turn to track the upper arm, which is similar to the head area. The CSK and IBCCF algorithms also drift to varying degrees. In the 44th frame, Iron Man continues to rotate, and the tracked facial area turns to the back. The CSK, CN2, DSST, and IBCCF algorithms all lose the target, and the Staple algorithm contains a small target area. At the 101st frame, an explosion occurs in the video and the image suddenly becomes brighter. At this time, only the algorithm in this chapter can accurately track the target.
[0247] Fig. 9 The accuracy and success rate curves of batch simulation of this algorithm and the comparative algorithm on the OTB100 dataset. Table 1 is the quantitative indicators of the accuracy and success rate obtained by simulating the algorithms in this chapter and the comparative algorithms. The first one is indicated by bold and underline, and the second one is indicated by bold.
[0248] Table 1 Performance indicators of each algorithm after data summary
[0249] JFAFMT CSK CN2 DSST IBCCF Staple AUC <![CDATA[ 0.808 ]]> 0.541 0.364 0.746 0.783 0.778 precision <![CDATA[ 0.598 ]]> 0.386 0.253 0.556 0.594 0.572 FPS 12.82 <![CDATA[ 240.26 ]]> 88.58 21.97 1.78 16.36
[0250] As shown in Table 1, in all the batch simulation results of OTB100, the CLE and OR curves of this algorithm are at the top. In the AUC ranking, the algorithm in this chapter ranks first with a score of 0.808, and the IBCCF algorithm ranks second with a score of 0.783. In the precision ranking, the algorithm in this chapter ranks first with a score of 0.598, and the IBCCF algorithm ranks second with a score of 0.594.
[0251] Table 2 shows the performance improvement of the algorithm in this chapter compared with other algorithms. Compared with CSK, CN2, DSST and Staple algorithms, the algorithm in this chapter has decreased in speed but increased in accuracy, with AUC increased by 49.35%, 121.98%, 8.31% and 3.86% respectively, and precision increased by 54.92%, 136.36%, 7.55% and 4.55% respectively. Compared with the IBCCF algorithm, all aspects have been improved, with AUC increased by 3.19%, precision increased by 0.67%, and FPS increased by 620.22%.
[0252] Table 2 Performance improvement rate of the algorithms in this chapter
[0253] JFAFMT CSK CN2 DSST IBCCF Staple AUC - 49.35% 121.98% 8.31% 3.19% 3.86% precision - 54.92% 136.36% 7.55% 0.67% 4.55% FPS - -94.66% -85.53% -41.65 620.22% -21.64%
[0254] In summary, when the target undergoes various types of changes, the algorithm proposed in this chapter has a higher accuracy rate than other algorithms and can cope with complex scenarios. The above is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A target tracking method based on multi-scale joint feature adaptive fusion, characterized in that: The method includes: Step 1, input the first frame image, and initialize the feature filter, scale filter and color histogram model; Step 2, input the t-th frame image, perform contour scale detection, extract and screen candidate frames, perform extreme scale scaling on the screened candidate frames to obtain the final candidate frames, and extract the candidate target area from the final candidate frames; Step 3, extract features for each candidate target region, and determine the HOG features, color histogram features, Haar local features, and LBP feature responses of each candidate target region; Step 4, calculating the adaptive fusion parameters, performing adaptive response fusion, and obtaining the final response result; determining the final target prediction candidate area according to the maximum value of the final response result, returning the candidate target area corresponding to the maximum response, and determining the scale of the target area as the target scale of the current frame; Step 5, using the final target prediction candidate region of the t-th frame image to construct a basic training sample, training and updating the filter model and the color histogram model, and obtaining the filter model and the color histogram model of the next frame image; Step 6: Repeat steps 2 to 5 until the tracking is completed.
2. The tracking method based on multi-scale joint feature adaptive fusion according to claim 1, characterized in that: The first frame image is input, and the feature filter, the scale filter and the color histogram model are initialized, specifically including: The target area of the first frame image has been manually marked. The target area is used as the basic sample and a training sample is constructed by cyclic shifting. Using the ridge regression method, the initialization filter model ω is obtained through the training samples. Where a represents the basis vector corresponding to the sample, b represents the standard response corresponding to the sample, ω represents the filter model parameter to be solved, and λ represents the regularization coefficient; The kernel function method is used to optimize the filter model ω, and the optimized solution of the filter model is obtained as follows: In the formula, Represents the calculation of the correlation between the sample and its own kernel, and I represents the identity matrix; Obtaining an initialized color histogram model using the target area as a basic sample; An initialized scale filter model is obtained using the target area as a basic sample.
3. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 1 or 2, characterized in that: The input t-th frame image is subjected to contour scale detection, candidate frames are extracted and screened, the screened candidate frames are subjected to extreme scale scaling to obtain a final candidate frame, and a candidate target area is extracted from the final candidate frame, specifically including: The contour of the target area of the t-1th frame image is extracted, and the edge calculation is performed on each pixel point to obtain the edge response value of each point, and then the contour sparse response map is obtained by the non-maximum suppression algorithm; each pixel point a in the contour sparse response map has a contour value size l a and contour direction β a , for each point's contour value l a Make a judgment, a If it is greater than 0.15, the point is considered as a contour and is retained; Contour clustering is performed on the contour sparse response map, and the contour responses of the 8-connected domains around each pixel in the target area are compared. If the contour directions β of two pixels are a If the difference is less than 90°, they are judged to be the same type of contour, otherwise they are not the same type of contour; The correlation of the contour clustering results is judged. For the center coordinates (x i ,y i ) of the contour class e i With the center coordinates (x j ,y j ) of the contour class e j ,pass Determine the correlation between two contour classes, E c (e i ,e j ) represents the correlation between two contour classes, β ij is the coordinate (x i ,y i ) and coordinates (x j ,y j ), α c Represents the relevant parameters, Δ d Represents the shortest distance between two classes; Determine each contour class e i The weight value ξ b , S eb Represents the candidate bounding box b wh With contour class i The area of the region included, S e Represents the contour class e i Area covered; According to each contour class e i The weight value ξ b Determine each contour class e i The final confidence p of the candidate bounding box b b ; Extract candidate bounding boxes and retain the top 35% of the candidate bounding boxes with confidence; Compare the area of the candidate bounding boxes with the top 35% confidence level with the target area determined by the t-1 frame image, and retain the candidate bounding boxes that are larger than 0.6 times and smaller than 1.4 times the target area of the t-1 frame image; The candidate target areas are extracted by scaling the candidate regions retained by the extreme scale. The center point of each candidate region calibrated by the candidate box is (x, y). The target width and height determined by the t-1 frame image are w and h respectively. Each candidate region is scaled to a certain scale to extract candidate target image regions of i scales.
4. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 3 is characterized in that: The input t-th frame image first extracts the gradient histogram (HOG) features of the candidate target area, including the following specific steps: The RGB image of the candidate area is converted into a grayscale image by the following formula; Perform Gamma correction on the grayscale converted image; By using [-1,0,1] as the horizontal gradient operator, [-1,0,1] T As a vertical gradient operator, determine the horizontal image gradient vector and the vertical image gradient vector of the corrected image; Determine the gradient amplitude |G(x,y)| and the gradient direction θ(x,y) according to the horizontal image gradient vector and the vertical gradient vector; The corrected image is divided into small cell units Cell l, and the gradient direction in each cell is divided into 9 parts from 0° to 180°. Every 20° direction is represented by a direction area. Z1 to Z9 represent 9 direction areas respectively. The gradients of the 9 parts are weighted statistically to obtain the gradient direction histogram of Cell l, that is, the 9-length feature vector; Several adjacent cells are synthesized and merged into a block. In each block, the feature vectors of all cells in it are connected in series to form the HOG feature of the current block. The gradient histogram in the block is normalized to obtain the normalized gradient amplitude. The HOG features in all blocks are integrated to obtain 36-dimensional HOG features.
5. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 4 is characterized in that: The step of determining the color histogram feature of each candidate target area specifically includes: According to the Bayesian statistical formula, the likelihood probability that pixel x belongs to the target is In the formula, O is the foreground area of the candidate target area, and its surrounding space is the background area S. I is the target candidate area composed of the foreground area O and the background area S. b represents the classification interval of the color histogram. x Represents the combination of RGB values of pixel x in interval b; The likelihood term is estimated for the likelihood probability. In the formula, Indicates the number of pixels in the b classification interval in the O region in I, represents the number of pixels in the b classification interval in the S region of I, |O| represents the area of the foreground area, and |S| represents the area of the background area; Convert the likelihood probability of pixel point x belonging to the target into Here, 0.5 represents the estimated response value.
6. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 5 is characterized in that: Determining the Haar local features of each candidate target area specifically includes: The Haar local features include edge Haar features, linear Haar features, center Haar features and diagonal Haar features; Constructing the integral graph Among them, the pixel value ii(x,y) at the position (x,y) of the candidate area is the sum of the grayscale values i(x',y') of all pixels in the upper left corner of the original image (x,y), s(x,y) represents the cumulative sum in the row direction, s(x,-1)=0 is initialized, ii(x,y) represents an integral image, ii(-1,y)=0 is initialized; Scan the image line by line, recursively calculate the cumulative sum s(x,y) of each pixel (x,y) in the row direction and the value of the integral image ii(x,y), which are s(x,y)=s(x,y-1)+i(x,y) and ii(x,y)=ii(x-1,y)+s(x,y) respectively; Traverse the candidate target area, and when the pixel in the lower right corner of the image is scanned, the image integral map ii (x, y) is constructed; Calculate the Haar eigenvalues in the candidate target area according to the integral image; The Haar features include 2 edge Haar features, 1 center Haar feature, 4 linear Haar features and 1 diagonal Haar feature; To calculate the edge Haar feature, it is necessary to search the integral image 6 times, to calculate the line feature, it is necessary to search 8 times, to calculate the center Haar feature, it is necessary to search 8 times, and to calculate the diagonal Haar feature, it is necessary to search 9 times. The feature extraction is performed in the form of a sliding window, sliding 1 pixel unit each time. For each sliding window, the 8 Haar features in the window are calculated.
7. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 6 is characterized in that: Determining the LBP features of each candidate target area specifically includes: Select a 3×3 neighborhood of each pixel in the candidate target area, take the gray value of the central pixel as the reference, and compare the gray values of the pixels in the surrounding 8 neighborhoods. If the value of the neighborhood pixel is greater than or equal to the value of the central pixel, mark the position as 1, otherwise it is marked as 0. The LBP feature calculation formula is expressed as In the formula, (x c ,y c ) represents the center pixel coordinates, p represents the pth pixel in the neighborhood, i c Represents the gray value of the center pixel, i p represents the gray value of the neighborhood pixel, s(·) is the sign function, expressed as 8. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 7 is characterized in that: The adaptive fusion parameters are calculated, adaptive response fusion is performed, and a final response result is obtained; According to the maximum value of the final response result, determining the final target prediction candidate area, specifically including: The four feature response weights were evaluated by the average peak correlation energy (APCE). In the formula, F max Indicates the maximum value in the response graph, F min represents the minimum value in the response graph, (w,h) represents the response coordinates; Step 4 obtains the response of each feature, which is recorded as r es1 、r es2 、r es3 、r es4 , calculate the APCE of each characteristic response as P APCE1 , P APCE2 , P APCE3 , P APCE4 ; Construct a mapping function to make the APCE ratio of each response float within a reasonable range. The mapping function is: In the formula, ω ss represents the pre-fusion parameter of the response, γ m As a hyperparameter, the parameter ω can be ss Keep it within a reasonable range; Determine the parameter ω of each target area image ss And normalize it, the obtained fusion parameter ω s for According to the fusion parameter ω s Determine the final fusion response of the four features of each target area as r es =ω s1 × es1 +ω s2 × es2 +ω s3 × es3 +ω s4 × es4 ; In the formula, r es Represents the final response result, returns the candidate target area corresponding to the maximum response result as the final target candidate area, and sets the scale of the target area (w i ,h i ) is determined as the target scale for the current frame.
9. The target tracking method based on multi-scale joint feature adaptive fusion according to claim 8, characterized in that: The method of constructing a basic training sample, training and updating the filter model and the color histogram model, and obtaining the filter model and the color histogram model of the next frame image specifically includes: The final target candidate area is used as the basic sample to construct the training sample. The training samples and the target candidate regions form a kernel matrix, K z =C(k az );where K z represents the kernel matrix composed of training samples and target candidate regions, k az Represents the kernel correlation calculation between the training sample a and the sample z in the area to be detected; Determine the objective function expression according to the kernel matrix Get the predicted response R(z) in time domain, R(z) = K z α; where K z represents the kernel matrix composed of training samples and target candidate regions, α represents the correlation filter model coefficient of the previous frame image, and the target position is the position corresponding to the maximum value of R(z); Update the color histogram model x and the coefficient α of the filter model, x 1:t =(1-η)x 1:t-1 +ηx t , α 1:t =(1-η)α 1:t-1 +ηα t ; In the formula, η represents the learning rate, x t represents the target sample vector of the current frame, x1:t represents the target sample vector learned from the beginning to time t, α t Represents the filter coefficient of the current frame, and α1:t represents the filter coefficient learned from the beginning to time t.
Citation Information
Patent Citations
Multi-characteristic fusion scale adaptive target tracking method
CN108510521A
Target tracking method, device and equipment and computer readable storage medium
CN110009663A
Scale self-adaptive sea surface target tracking method based on edge detection
CN110111369A
Target tracking method and system based on hierarchical convolution characteristics and scale self-adaptive kernel correlation filtering
CN110120065A
Long-time video tracking method based on adaptive correlation filtering
CN110472577A