A video object tracking method
Through the combination of Bayesian classifier and sparse feature template, the accuracy and stability problems of the video target tracking algorithm under lighting and deformation are solved, and stable target tracking is achieved.
Patent Information
- Application Number
- CN202211358869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-11-01
AI Technical Summary
The tracking accuracy and stability of existing video target tracking algorithms in the case of target lighting and deformation are reduced, and even the target may be lost.
A video target tracking method is adopted, using Bayesian classifier combined with sparse feature templates, randomly generate positive and negative sample image blocks, calculate feature vectors, and classify candidate image blocks in the search box, and update Bayesian classifier parameters to adapt to target changes.
The target tracking accuracy and stability in the case of light changes and deformation is significantly improved, the target loss is avoided, and the stable target position output is achieved.
Smart Images

Figure CN115661205B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of video image processing, and particularly relates to a video target tracking method, which can be used for video target detection and tracking, moving target analysis, etc. Background Art
[0002] With the continuous development of the field of digital image processing, video target tracking technology, as a key step in the target analysis process, is also constantly developing and progressing. Video target tracking refers to the process of positioning the same target in each frame of the input video containing the target area and outputting the latest position information of the target as the tracking result of the current frame.
[0003] Currently, the most widely used video target tracking method is the correlation tracking method. The correlation tracking method first inputs the target video, and at the same time, manually gives the target gate (including target position and size information) to generate a target template. Then, according to a certain optimization criterion, it searches for the latest position where the target template appears in the subsequent video image frames and outputs this position as the target tracking result. This method has a simple principle and is easy to implement. However, if the target in the current image is affected by illumination changes or deformation, the shape of the target image area will become incomplete or produce large deformations, which will affect the accuracy of the target template position search and even cause the failure to find the target template. Therefore, the tracking accuracy and stability of the existing tracking algorithms will decrease in the case of target illumination and deformation, and even the target will be lost. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem that the tracking accuracy and stability of the existing tracking algorithms will decrease in the case of target illumination and deformation, and even the target will be lost, and to propose a video target tracking method to reduce the influence of target illumination and deformation on the stability of target tracking accuracy.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A video target tracking method, characterized in that it includes the following steps:
[0007] Step 1: Read a frame of the video image to obtain image M and the target in image M;
[0008] Step 2: Based on the target in image M, set the target gate. If image M is the first frame image of the video, execute Step 3; otherwise, execute Step 6;
[0009] Step 3: Randomly generate e positive sample image blocks and d negative sample image blocks within the target gate, where e≥2, d≥2, and there is no overlapping area between the positive sample image blocks and the negative sample image blocks;
[0010] Step 4: Based on the image M and the target gate, calculate the feature values of each positive sample image patch and negative sample image patch, and obtain the template feature vector;
[0011] Step 5: Initialize the Bayesian classifier H(v) based on the template feature vector in Step 4, and return to Step 1;
[0012] Step 6: Draw a search box, select R candidate image patches within the search box, calculate the template feature vector of each candidate image patch, classify using the Bayesian classifier, and obtain and output the target position in the next frame image; the candidate image patches refer to the possible positions of the target in the next frame image, where R≥2;
[0013] If the target position meets the tracking requirements, end the tracking; otherwise, execute Step 7;
[0014] Step 7: Calculate the template feature vector corresponding to the target position in Step 6, update the parameters of the Bayesian classifier H(v), and then return to Step 1.
[0015] Furthermore, in Step 3, the method of randomly generating e positive sample image patches and d negative sample image patches is as follows:
[0016] Definition: The target gate is (t x ,t y ,t w ,t h );
[0017] The extraction range of positive sample image patches is a rectangular area centered at the coordinate (t x ,t y ), with [p w ,p h as the width and height, extending to the left and right and up and down respectively to form a rectangular area. Randomly generate e positive sample image patches within the rectangular area, where 2p w <t w ,2p h <t h ;
[0018] The extraction range of negative sample image patches is a rectangular area outside a rectangular area centered at (t x ,t y ) with a width of 2nwo and a height of 2nho, and inside a rectangular area with a width of 2nwi and a height of 2nhi. Randomly generate d negative sample image patches within the rectangular area; where nwo≥p w , nho≥p h , 2nwi≤t w ,2nhi≤t h .
[0019] Further, step 4 specifically includes the following steps:
[0020] 4.1 Generate a sparse feature template based on the target gate.
[0021] 4.2 Set the sparse feature template at T different positions of the target gate. Meanwhile, randomly generate 2 to 4 rectangular blocks at each position, and obtain the four corner coordinates of each rectangular block and randomly generate the weight of each rectangular block, where T ≥ 2.
[0022] Randomly generate the weight of each rectangular block through the following formula:
[0023]
[0024] Where randi(2) is a random number operator with an output value of 1 or 2; Q r represents the number of randomly generated rectangular blocks when the sparse feature template is at the r-th position of the target gate. wrl represents the weight of the l-th rectangular block when the sparse feature template is at the r-th position of the target gate. 2 ≤ Q r ≤ 4, 1 ≤ l ≤ Q r , 1 ≤ r ≤ T;
[0025] 4.3 Based on the image M in step 1, obtain the integral image H of the grayscale image of image M, and calculate the feature values of each positive sample image block and each negative sample image block based on the integral image H, the weight of each rectangular block in step 4.2, and the four corner coordinates, and obtain the template feature vectors of each positive sample image block and each negative sample image block.
[0026] Further, step 4.3 specifically includes the following steps:
[0027] 4.3.1 Based on the image M in step 1, obtain the integral image H of image M.
[0028] 4.3.2 Calculate the feature values of each positive sample image block and each negative sample image block based on the integral image H, the weight of each rectangular block in step 4.2, and the four corner coordinates of the corresponding rectangular block, and obtain the template feature vectors of each positive sample image block and each negative sample image block.
[0029] 4.3.2.1 Calculate the feature values of each positive sample image block and each negative sample image block based on the integral image H, the weight of each rectangular block in step 4.2, and the four corner coordinates of the corresponding rectangular block.
[0030] When the sparse feature template is at the r-th position of the target gate, the linear sum of the four corner coordinates of the -th positive sample image block and the four corner coordinates of the corresponding position of the l-th rectangular block is Among them, set the upper left corner of the image M as the origin coordinates, with the right direction as the X-axis and the downward direction as the Y-axis to establish a two-dimensional coordinate system. The linear sum of the angular coordinates of the th positive sample image block close to the origin and the corresponding angular coordinates of the lth rectangular block is respectively the linear sums of the angular coordinates in the counterclockwise direction;
[0031] Then the eigenvalue of the lth rectangular block is calculated by the following formula:
[0032] where H(·) is the integral image value operator;
[0033] Based on each negative sample image block, the method for calculating the eigenvalue of the lth rectangular block is the same as that based on each positive sample image block;
[0034] 4.3.2.2 Obtain the eigenvalues of each positive sample image block and each negative sample image block;
[0035] The eigenvalue of the th positive sample image block is calculated by the following formula:
[0036]
[0037] The method for obtaining the eigenvalue of each negative sample image block is the same as that for obtaining the eigenvalue of each positive sample image block;
[0038] 4.3.2.3 Obtain the template feature vectors of each positive sample image block and each negative sample image block.
[0039] Furthermore, the specific steps of step 5 are as follows:
[0040] 5.1 Based on the template feature vectors in step 4, obtain the mean and variance of the positive sample conditional probability and the mean and variance of the negative sample conditional probability;
[0041] 5.2 Based on the obtained mean and variance of the positive sample conditional probability and the mean and variance of the negative sample conditional probability, initialize the parameters of the Bayesian classifier H(v).
[0042] Furthermore, the specific steps of step 6 are as follows:
[0043] 6.1 Draw a search box, select R candidate image blocks within the search box, and calculate the feature vectors of each candidate image block using the same methods as in steps 2, 3, and 4;
[0044] 6.2 Obtain and output the target position of the next frame of the image;
[0045] 6.2.1 Based on the feature vectors of each candidate image patch, obtain the variance and mean of each candidate image patch's feature vector, and calculate the probability that each candidate image patch belongs to a positive sample image patch and the probability that it belongs to a negative sample image patch. The calculation formulas are as follows:
[0046]
[0047] Among them, p1 b represents the probability that the b-th candidate image patch belongs to a positive sample image patch, and p0 b represents the probability that the b-th candidate image patch belongs to a negative sample image patch. is the parameter of the Bayesian classifier H(v) in the current state, where 1 ≤ b ≤ R; v b is the b-th candidate image patch;
[0048] 6.2.2 Calculate the probability difference pd b = p1 b - p0 b of each candidate image patch, and the candidate image patch corresponding to the maximum probability difference is the target position, and output the target position; b
[0049] 6.3 If the target position meets the tracking requirements, end the tracking; otherwise, execute step 7.
[0050] Furthermore, in step 7, use the same method as in steps 2, 3, 4, and 5 to calculate the template feature vector corresponding to the target position in step 6, and update the parameters of the Bayesian classifier H(v) according to the template feature vector corresponding to the target position in step 6.
[0051] Furthermore, in steps 6 and 3, the weights of the rectangular blocks corresponding to the target gating are kept consistent.
[0052] Furthermore, in step 7, the formula for calculating the parameters of the Bayesian classifier H(v) is as follows:
[0053]
[0054] Among them, represents the updated parameter of the Bayesian classifier H(v), represents the parameter of the Bayesian classifier H(v) in the current state, represents the mean and variance of the image composed of all positive sample image patches among the feature vectors corresponding to the T image patch groups at the target position, represents the mean and variance of the image composed of all negative sample image patches among the feature vectors corresponding to the T image patch groups at the target position, and λ represents the adjustment variable of the parameter update speed, where 0 ≤ λ ≤ 1.
[0055] The beneficial effects of the present invention are as follows:
[0056] 1. The present invention uses a video object tracking method based on sparse features to stably extract and represent the positive sample features and negative sample features of an object affected by illumination changes or deformations, and uses a Bayesian classifier to search for the object in the current image, determine the latest position of the object and output it.
[0057] Compared with the related tracking methods currently adopted, the present invention can significantly improve the object tracking accuracy and stability under illumination changes and deformations.
[0058] 2. In order to further improve the stable tracking effect of the object under illumination changes and deformations, this method updates the feature vector of the object in each object tracking session to adapt to the changes that occur to the object itself during the tracking process, avoid object loss, and avoid the occurrence of tracking failure caused by gradually changing illumination during the tracking process. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flowchart of an embodiment of the present invention;
[0060] Figure 2 is a schematic diagram of the calculation of the integral image H in an embodiment of the present invention;
[0061] Figure 3 is an example of the calculation steps of the integral image H in an embodiment of the present invention;
[0062] Figure 4 is an example of a sparse feature template in an embodiment of the present invention;
[0063] Figure 5 is a schematic diagram of the positive sample extraction area in an embodiment of the present invention;
[0064] Figure 6 is a schematic diagram of the negative sample extraction area in an embodiment of the present invention;
[0065] Figure 7 is a schematic diagram of the eigenvalues of three rectangular blocks in an embodiment of the present invention;
[0066] Figure 8 is a schematic diagram of the search area of the candidate image block in an embodiment of the present invention;
[0067] Figure 9 Shows the tracking effect of the aircraft target video in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0068] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. Specifically, an input video with a target and a spatial resolution of W*H (width * height) will be used to elaborate on the present invention in detail.
[0069] The present invention proposes a video target tracking method. As Figure 1 shown, taking one target as an example, it includes the following steps:
[0070] Step 1: Read a frame of the input video and obtain image M, where image M includes 1 target; the input video can be a video file stored on a computer hard disk, and the video spatial resolution of the video file is W*H;
[0071] Step 2: Based on the target in image M, set a target gate. If image M is the first frame image of the video, execute Step 3; otherwise, execute Step 6;
[0072] Step 3: Randomly generate e positive sample image patches and d negative sample image patches within the target gate, where e≥2, d≥2, and there is no overlapping area between the positive sample image patches and the negative sample image patches;
[0073] Specifically: The method for obtaining positive sample image patches and negative sample image patches is as follows:
[0074] As Figure 5 shown, define the target gate as (t x , t y , t w , t h ), (t x , t y ) is the coordinate value of the center point of the target gate (the coordinate value of the center point of the target gate is the coordinate value of any point of the target), t w is the width of the target, t h is the height of the target, that is, the target gate (t x , t y , t w , t h ) is a rectangular area centered on (t x , t y ) with t w and t h as the side lengths;
[0075] In the target gate, the extraction range of the positive sample image patch is a rectangular area formed by extending left and right and up and down with (t x , t y ) as the center and [p w , p h as the width and height. As Figure 5 shown, that is, it encloses a rectangle centered on (t x , ty ) centered, with a width of 2p w , and a height of 2p h rectangular region, and randomly generate e positive sample image patches within the rectangular region, where 2p w < t w , 2p h < t h ;
[0076] As Figure 6 shown, in the target gate, the internal parameters of the extraction range of the negative sample image patches are [nwo, nho], the external parameters are [nwi, nhi], and the center point coordinate value is (t x , t y ), then the extraction region of the negative sample image patches is an annular region outside a rectangular region centered at (t x , t y ) with a width of 2nwo and a height of 2nho, and inside a rectangular region with a width of 2nwi and a height of 2nhi. Randomly generate d negative sample image patches within the annular region; where nwo ≥ p w , nho ≥ p h , nwo ≥ p w and nho ≥ p h The purpose is to ensure that there is no overlapping region between the negative sample image patches and the positive sample image patches. In practical applications, nwi = 3nwo, nhi = 3nho, and e < d; where 2nwi ≤ t w , 2nhi ≤ t h ;
[0077] Step 4: Based on the image M and the target gate, calculate the feature values of each positive sample image patch and each negative sample image patch, and obtain the template feature vector;
[0078] Specifically:
[0079] 4.1 As Figure 4 shown, based on the target gate, generate a sparse feature template. The sparse feature template is a rectangle formed by the t w and t h of the target;
[0080] 4.2 Set the sparse feature template at T different positions of the target gate, and at the same time randomly generate 2 - 4 rectangular blocks at each position, and obtain the four corner coordinates of each rectangular block and randomly generate the weight of each rectangular block, where T ≥ 2;
[0081] The main function of the rectangular block is to provide data for calculating the feature values of the positive sample image patches and the negative sample image patches, and use the rectangular block to efficiently describe the target appearance features;
[0082] The generation method of the rectangular block is: AsFigure 4 As shown, when the sparse feature template is in different positions, 2 to 4 rectangular blocks are randomly generated within the sparse feature template. The position, size, and weight of each rectangular block are also randomly generated. The method for generating the weight of each rectangular block is as follows:
[0083]
[0084] Among them, randi(2) is a random number operator with an output value of 1 or 2; Q r represents the number of randomly generated rectangular blocks when the sparse feature template is in the r-th position of the target gate. wrl represents the weight of the l-th rectangular block when the sparse feature template is in the r-th position of the target gate. 2 ≤ Q r ≤ 4, 1 ≤ l ≤ Q r , 1 ≤ r ≤ T;
[0085] For the convenience of calculation, it is necessary to ensure that: each rectangular block has two of its sides parallel to the X-axis; in this embodiment, T = 50;
[0086] In this embodiment, the described rectangular blocks can be generated offline and remain unchanged during the target tracking process, that is, the weight of each rectangular block can remain unchanged;
[0087] Set the upper left corner of image M as the origin coordinate, with the right direction as the X-axis and the downward direction as the Y-axis to establish a two-dimensional coordinate system. Then, according to the two-dimensional coordinate system, the four corner coordinates of each rectangular block are obtained;
[0088] 4.3 Based on the image M in step 1, obtain the grayscale image of image M, and further obtain the integral image H of the grayscale image. Based on the integral image H, the weight of each rectangular block in step 4.2, and the four corner coordinates of the corresponding rectangular block, calculate the feature values of each positive sample image block and negative sample image block, and obtain the template feature vectors of each positive sample image block and negative sample image block;
[0089] 4.3.1 Based on the image M in step 1, obtain the grayscale image of image M, and further obtain the integral image H of the grayscale image;
[0090] The specific method is as follows:
[0091] The main function of calculating the integral image H is to provide data for calculating the feature vectors of positive sample image blocks and negative sample image blocks;
[0092] Such as Figure 2 And Figure 3As shown in the figure, for the grayscale image of image M, the pixel value at any point coordinate (B, C) in its integral image H is the sum of the grayscale values of all pixel points within the rectangular area formed from the origin (0, 0) of the grayscale image to the point (B, C). Its calculation formula is:
[0093] where Z = 0…B, V = 0…C
[0094] As Figure 2 and Figure 3 shown, first, traverse the pixel grayscale values of each row of the grayscale image one by one to the right along the positive X-axis, and replace the current pixel grayscale value with the sum of the current pixel value and the pixel value at the previous position (the pixel value of the first pixel from the left in each row remains unchanged). After traversing all pixel values along the positive X-axis, obtain the intermediate image I. Then, traverse the pixel grayscale values of each column of image I one by one downward along the negative Y-axis, and replace the current pixel grayscale value with the sum of the current pixel value and the pixel value at the previous position (the pixel value of the first pixel from the top in each column remains unchanged). After traversing all pixel values, obtain the integral image H;
[0095] 4.3.2 Based on the integral image H, the weights of each rectangular block in step 4.2, and the four corner coordinates of the corresponding rectangular block, calculate the feature values of each positive sample image block and each negative sample image block, and obtain the template feature vectors of each positive sample image block and each negative sample image block;
[0096] Specifically: First, select any one position from the 50 neighborhood positions of the target gating. Assuming that the sparse feature template is at this position, the randomly generated number of rectangular blocks Q1 = 3, that is, the sparse feature template randomly generates 3 rectangular blocks with different sizes, positions, and weights at this position. The calculation method for the feature value of the image block is the weighted sum of all rectangular blocks inside the sparse feature template and the feature values of the corresponding image blocks, where the feature value of each rectangular block is the integral image value linearly calculated from the four corner coordinates of this rectangular block and the corresponding four corner coordinates of the corresponding image block;
[0097] The calculation method is as follows:
[0098] Taking one positive sample image block as an example:
[0099] As Figure 7 shown, in the coordinate axis, (that is, ) respectively represent that at the r-th position, the four corner coordinates of the first rectangular block and the corresponding four corner coordinates of the th positive sample image block are linearly summed, respectively represent the integral image values of the linear sums of the four corner coordinates, (that is, ) respectively represent that at the r-th position, the four corner coordinates of the second rectangular block are the linear sum of the corresponding four corner coordinates of the -th positive sample image block, respectively represent the integral image values of the linear sums of the four corner coordinates, (i.e., ) respectively represent that at the r-th position, the four corner coordinates of the third rectangular block are the linear sum of the corresponding four corner coordinates of the -th positive sample image block, respectively represent the integral image values of the linear sums of the four corner coordinates; then, the eigenvalue calculation formula for each rectangular block is as follows:
[0100]
[0101] where, H(·) is the operator for taking the integral image value;
[0102] then for the -th positive sample image block, the eigenvalue calculation formula when the corresponding sparse feature template is at the r-th position is as follows:
[0103] where, wr1, wr2, wr3 are the weights of the three rectangular blocks respectively;
[0104] Obtain the template feature vector of the -th positive sample image block, specifically:
[0105] represents the feature vector of the -th positive sample image block, which is a 50-dimensional vector;
[0106] Calculate the template feature vector of each negative sample image block using the same method;
[0107] Step 5: Based on the template feature vectors of each positive sample image block and each negative sample image block in Step 4, initialize the Bayesian classifier H(v), and return to Step 1;
[0108] 5.1 Based on the 50 groups of template feature vectors calculated in Step 4, obtain the mean and variance of the positive sample conditional probability and the mean and variance of the negative sample conditional probability;
[0109] Specifically: Since the template feature vectors approximately follow a Gaussian distribution, it can be considered that in the Bayesian classifier H(v), the conditional probability p(v r |y = 1), r = 1...50 of the positive sample image block and the conditional probability p(v r |y = 0) of the negative sample image block approximately follow a Gaussian distribution where v r The sparse feature template at the r-th position of the table;
[0110] Set the initial parameters of the Bayesian classifier H(v) as Before initialization, calculate the mean and variance of the positive sample conditional probability r = 1...50, and the mean and variance of the negative sample conditional probability r = 1...50;
[0111] 5.2 Based on obtaining the mean and variance of the positive sample conditional probability r = 1...50 and the mean and variance of the negative sample conditional probability r = 1...50, initialize the parameters of the Bayesian classifier H(v), that is, set the parameters of the Bayesian classifier H(v) Iterate to That is, let
[0112] Step 6: Draw a search box, select and set R candidate image patches within the search box, calculate the template feature vector of each candidate image patch, classify using the Bayesian classifier, and obtain and output the target position of the next frame of the image; the candidate image patches refer to the possible positions of the target in the next frame of the image, R≥2;
[0113] If the target position meets the tracking requirements, end the tracking; otherwise, execute Step 7;
[0114] 6.1 Draw a search box, select and set R candidate image patches within the search box, and calculate the template feature vector of each candidate image patch using the same method as in Steps 2, 3, and 4;
[0115] As Figure 8 shown, specifically: The method for extracting the image patches in the search window is: taking the target gate (t x ,t y ,t w ,t h ) as an example:
[0116] Set the range of the search window to (srw, srh), then the horizontal coordinate range of the candidate image patches is [t x -srw, t x +srw], and the vertical coordinate range is [t y -srh, t y +srh], as Figure 8 shown, take all the image patches in this rectangular area as candidate image patches, and set that there are R image patches in this rectangular area;
[0117] 6.2 Obtain and output the target position of the next frame image;
[0118] 6.2.1 Based on the template feature vectors of each candidate image patch, obtain the variance and mean of the template feature vectors of each candidate image patch, and calculate the probability that each candidate image patch belongs to a positive sample image patch and the probability that it belongs to a negative sample image patch;
[0119] The calculation formula is as follows:
[0120]
[0121] Among them, p1 b represents the probability that the b-th candidate image patch belongs to a positive sample image patch, and p0 b represents the probability that the b-th candidate image patch belongs to a negative sample image patch. is the parameter of the Bayesian classifier H(v) in the current state, 1 ≤ b ≤ k; v b is the b-th candidate image patch;
[0122] 6.2.2 Calculate the probability difference of each candidate image patch through the formula pd b = p1 b - p0 b The candidate image patch corresponding to the largest probability difference is the target position, and output the target position;
[0123] Step 7: Calculate the template feature vector corresponding to the target position in Step 6 using the same method as in Step 2, Step 3, Step 4, and Step 5, and update the parameters of the Bayesian classifier H(v) according to the template feature vector corresponding to the target position in Step 6, then return to Step 1;
[0124] The formula for calculating the parameters of the Bayesian classifier H(v) is as follows:
[0125]
[0126] Among them, represents the updated parameter of the Bayesian classifier H(v), represents the parameter of the Bayesian classifier H(v) in the current state, represents the template feature vector corresponding to the target position, the mean and variance of the image composed of all positive sample image patches represents the mean and variance of the image composed of all negative sample image patches among the feature vectors corresponding to the T image patch groups at the target position, and λ represents the adjustment variable of the parameter update speed, usually taking 0 ≤ λ ≤ 1, the smaller λ is, the faster the update speed;
[0127] After calculation, let After the iteration is completed, return to Step 1.
[0128] Figure 9 The figure shows the tracking effect of the target stable tracking algorithm adopting the solution of the present invention on the aircraft target video, and the effect images of the 1st, 37th, 93rd, 170th, 214th, and 414th frames are selected; it can be seen that the method in this article can accurately extract the target position under different illumination conditions of the target. For example, the target in the first frame is in the front light, and the targets in the 37th and 93rd frames are in the backlight conditions. Moreover, during the flight of the aircraft target, its shape changes continuously. As can be seen from the 6 effect images, the present invention can stably track it without losing the target, and the extracted target position is also relatively accurate.
[0129] In other embodiments of the present invention, the number of targets can be greater than 1, and the number of positive sample image blocks and negative sample image blocks within each target gate can be different.
[0130] In other embodiments of this solution, the region of the positive sample image block is not centered at (t x ,t y ), as long as it satisfies that there is no overlapping region between the positive sample image block and the negative sample image block, and both the positive sample image block and the negative sample image block are located within the target gate.
[0131] Table 1 gives Figure 9 the time consumption of the target tracking simulation operation for each image of (simulation conditions: CPU main frequency 3.6 GHz, memory 8 GB, Windows 7 operating system, programming language MATLAB, video image resolution 1024*1024, gray bit depth 8 bits);
[0132] Table 1
[0133] Frame number 1 37 93 170 214 414 Running time (milliseconds) 38 39 39 37 38 39
[0134] It can be seen that the target stable tracking algorithm in this document can meet the real-time target tracking task of 25 frames per second (that is, the algorithm running time for each frame of image is less than 40 milliseconds) under the above simulation conditions.
Claims
1. A video object tracking method, characterized in that, It includes the following steps: Step 1: Read a frame image of the video to obtain image M and the target in image M; Step 2: Based on the target in image M, set the target gate. If image M is the first frame image of the video, execute Step 3; otherwise, execute Step 6; Step 3: Randomly generate e positive sample image blocks and d negative sample image blocks within the target gate, where e≥2, d≥2, and there is no overlapping area between the positive sample image blocks and the negative sample image blocks; Step 4: Based on image M and the target gate, calculate the feature values of each positive sample image block and each negative sample image block, and obtain the template feature vector; Step 5: Initialize the Bayesian classifier H(v) based on the template feature vector in Step 4, and return to Step 1; Step 6: Draw a search box, select R candidate image blocks within the search box, calculate the template feature vector of each candidate image block, classify using the Bayesian classifier, and obtain and output the target position of the next frame image; the candidate image blocks refer to the possible positions of the target in the next frame image, R≥2; If the target position meets the tracking requirements, end the tracking; otherwise, execute Step 7; Step 7: Calculate the template feature vector corresponding to the target position in Step 6, update the parameters of the Bayesian classifier H(v), and then return to Step 1; The specific content of Step 4 includes the following steps: 4.1 Generate a sparse feature template based on the target gate; 4.2 Set the sparse feature template at T different positions of the target gate, and at the same time randomly generate 2 to 4 rectangular blocks at each position, and obtain the four corner coordinates of each rectangular block and randomly generate the weight of each rectangular block, where T≥2; Randomly generate the weight of each rectangular block through the following formula: Among them, randi(2) is a random number operator with an output value of 1 or 2; Q r represents the number of randomly generated rectangular blocks when the sparse feature template is in the r-th position of the target gate. wrl represents the weight of the l-th rectangular block when the sparse feature template is in the r-th position of the target gate. 2 ≤ Q r ≤ 4, 1 ≤ l ≤ Q r , 1 ≤ r ≤ T; 4.3 Based on image M in Step 1, obtain the integral image H of the grayscale image of image M, and based on the integral image H, the weight of each rectangular block in Step 4.2, and the four corner coordinates, calculate the feature values of each positive sample image block and each negative sample image block, and obtain the template feature vector of each positive sample image block and each negative sample image block.
2. A video target tracking method according to claim 1, characterized in that: In Step 3, the method of randomly generating e positive sample image blocks and d negative sample image blocks is as follows: Definition: The target gate is (t x , t y , t w , t h ); The extraction range of the positive sample image patches is a rectangular area formed by extending to the left and right and up and down with the center at the coordinates (t x ,t y ), with [p w ,p h as the width and height. e positive sample image patches are randomly generated within the rectangular area, where 2p w < t w , 2p h < t h ; The extraction range of negative sample image patches is an annular region outside a rectangular region centered at (t x , t y ) with a width of 2nwo and a height of 2nho, and inside a rectangular region with a width of 2nwi and a height of 2nhi. d negative sample image patches are randomly generated within the annular region; where nwo ≥ p w , nho ≥ p h , 2nwi ≤ t w , 2nhi ≤ t h .
3. A video target tracking method according to claim 2, characterized in that: The specific content of Step 4.3 includes the following steps: 4.3.1 Based on image M in Step 1, obtain the integral image H of image M; 4.3.2 Based on the integral image H, the weight of each rectangular block in Step 4.2, and the four corner coordinates of the corresponding rectangular block, calculate the feature values of each positive sample image block and each negative sample image block, and obtain the template feature vector of each positive sample image block and each negative sample image block; 4.3.2.1 Based on the integral image H, the weight of each rectangular block in Step 4.2, and the four corner coordinates of the corresponding rectangular block, calculate the feature values of each positive sample image block and each negative sample image block; When the sparse feature template is at the r-th position of the target gate, the linear sum of the four corner coordinates of the th positive sample image block and the four corner coordinates at the corresponding positions of the l-th rectangular block is where set the upper left corner of image M as the origin coordinate, with the right direction as the X-axis and the downward direction as the Y-axis to establish a two-dimensional coordinate system; Then the eigenvalue of the l-th rectangular block is calculated by the following formula: Among them, H(·) is an operator for obtaining integral image values; The method for calculating the eigenvalue of the l-th rectangular block based on each negative sample image patch is the same as that based on each positive sample image patch; 4.3.2.2 Obtain the eigenvalues of each positive sample image patch and each negative sample image patch; The eigenvalue of the positive sample image patch is calculated by the following formula: The method for obtaining the eigenvalue of each negative sample image patch is the same as that for obtaining the eigenvalue of each positive sample image patch; 4.3.2.3 Obtain the template feature vectors of each positive sample image patch and each negative sample image patch.
4. A video object tracking method according to claim 3, wherein: The specific steps of step 5 are as follows: 5.1 Based on the template feature vectors in step 4, obtain the mean and variance of the positive sample conditional probability and the mean and variance of the negative sample conditional probability; 5.2 Based on the obtained mean and variance of the positive sample conditional probability and the mean and variance of the negative sample conditional probability, initialize the parameters of the Bayesian classifier H(v).
5. A video object tracking method according to claim 4, wherein: The specific steps of step 6 are as follows: 6.1 Draw a search box, select R candidate image patches within the search box, and calculate the feature vector of each candidate image patch using the same method as in steps 2, 3, and 4; 6.2 Obtain and output the target position of the next frame of the image; 6.2.1 Based on the feature vector of each candidate image patch, obtain the variance and mean of the feature vector of each candidate image patch, and calculate the probability that each candidate image patch belongs to a positive sample image patch and the probability of a negative sample image patch. The calculation formula is as follows: Among them, p1 b represents the probability that the b-th candidate image patch belongs to a positive sample image patch, and p0 b represents the probability that the b-th candidate image patch belongs to a negative sample image patch, is the parameter of the Bayesian classifier H(v) in the current state, where 1 ≤ b ≤ R; v b is the b-th candidate image patch; 6.2.2 Calculate the probability difference pd of each candidate image patch through the formula b = p1 b - p0 b Calculate the probability difference pd of each candidate image patch b The candidate image patch corresponding to the maximum probability difference is the target position, and output the target position; 6.3 If the target position meets the tracking requirements, end the tracking; otherwise, execute step 7.
6. A video object tracking method according to claim 5, wherein: In step 7, calculate the template feature vector corresponding to the target position in step 6 using the same method as in steps 2, 3, 4, and 5, and update the parameters of the Bayesian classifier H(v) according to the template feature vector corresponding to the target position in step 6.
7. A video object tracking method according to claim 6, wherein: In steps 6 and 3, the weights of the rectangular blocks corresponding to the target gate are kept consistent.
8. A video object tracking method according to claim 7, wherein: In step 7, the formula for calculating the parameters of the Bayesian classifier H(v) is as follows: Among them, represents the updated parameters of the Bayesian classifier H(v), represents the parameters of the Bayesian classifier H(v) in the current state, represents the mean and variance of the image composed of all positive sample image patches among the feature vectors corresponding to the T image patch groups at the target position, represents the mean and variance of the image composed of all negative sample image patches among the feature vectors corresponding to the T image patch groups at the target position, and λ represents the adjustment variable of the parameter update speed, where 0 ≤ λ ≤ 1.
Citation Information
Patent Citations
Video target tracking method based on compressive sensing
CN104392467A
Accurate target tracking method on condition of severe shielding
CN108549905A