A target tracking method, device and medium based on asymmetric background perception
Through a method based on asymmetric background perception, asymmetric background perception matrix is calculated and filters introduced with spatial mask constraints are combined with depth features and channel attention weights, the accuracy and robustness of target tracking in the prior art under the background chaos, occlusion, scale changes and intra-class interference are solved, and more efficient target tracking is achieved.
Patent Information
- Application Number
- CN202311043712.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-08-18
AI Technical Summary
Existing related filter tracking algorithms are difficult to maintain high accuracy and robustness in the case of background messy, occlusion, scale changes and intra-class disturbances, and the depth characteristics are not enough to distinguish between targets and intra-class disturbances.
A target tracking method based on asymmetric background perception is proposed. By calculating the asymmetric background perception matrix and introducing filters with spatial mask constraints, a multimodal target pool is constructed to evaluate the reliability of candidate samples, and combined with deep feature extraction and channel attention weights, the discriminant ability and anti-interference ability of tracking are enhanced.
Improve the accuracy and robustness of target tracking, enhance the ability to distinguish background and target, and effectively solve the challenges brought by occlusion, scale changes and intra-class interference.
Smart Images

Figure CN116993782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target tracking, and in particular to a target tracking method, device and medium based on asymmetric background perception. Background Art
[0002] Visual target tracking technology aims to locate the position and size of the target in subsequent video frames by modeling the appearance features of the target given the position and size of the target in the first frame. It is widely used in the fields of drone target positioning, human-computer interaction, military target tracking, etc. However, target tracking is a very challenging basic problem. On the one hand, it is easily disturbed by external factors such as background clutter, occlusion, scale changes, etc., and on the other hand, it is restricted by internal factors such as limited computing power and power capacity of the equipment.
[0003] In recent years, the correlation filter (CF) tracking method has attracted widespread attention from scholars for its precise tracking accuracy and real-time tracking speed. The basic idea of the CF tracker is: first design an initial filter based on the first frame image, then perform a correlation operation between the filter and the next frame image, estimate the position and scale of the target in the next frame image based on the position and scale of the maximum filter response, and finally update the filter based on the appearance model of the target in the next frame. The method based on the Discriminative Correlation Filters (DCF) converts the correlation operation in the spatial domain into a dot product operation in the frequency domain, and greatly improves the robustness and accuracy at a satisfactory speed, becoming the mainstream method in the field of visual tracking.
[0004] Existing correlation filter tracking algorithms have several limitations: First, traditional correlation filter tracking methods only use symmetrical sampling to extract image features, while regular sampling ignores the shape information of the target, resulting in insufficient feature discrimination ability; second, traditional correlation filter tracking methods only focus on the difference between the background and the target and ignore the more challenging influence of intra-class interference; third, the correlation filter algorithm does not further evaluate the reliability of the best candidate samples, which makes it easy to fail in occluded scenes.
[0005] The uniqueness of the DCF method lies in converting the cyclic correlation calculation in the spatial domain into point-by-point element multiplication in the frequency domain through fast Fourier transform, thereby avoiding the inversion operation of large matrices and greatly improving the computational efficiency. The Fourier transform is based on periodic boundary conditions. This assumption causes the negative samples generated by the cyclic shift to have boundary effects that are inconsistent with the actual scenario, resulting in a decrease in the generalization ability of the model.
[0006] Deep learning technology has greatly advanced the task of visual object tracking by providing powerful deep feature learning capabilities. Siamese network-based trackers formulate the visual object tracking problem as a matching problem by calculating the mutual correlation similarity between the target template and the search area. The latest work attempts to combine deep features with traditional correlation tracking algorithms, such as MDNet, C-COT, ECO, and GFS-DCF. Thanks to the high-level semantic features extracted by deep learning, such methods effectively improve the target and background discrimination ability of the tracker. This is because the target and background are essentially inter-class objects with different semantic characteristics. It should be pointed out that if there are highly similar intra-class distractors in the video frame, the deep features are not enough to distinguish between the target and the distractors. In this case, the correlation filter is easily misled by the distractors, resulting in filter degradation.
[0007] Due to limitations in feature recognition capabilities, computing power and power consumption, as well as occlusion, severe deformation, and interference from similar objects, designing a real-time, high-precision, and high-robust tracker remains a very challenging scientific problem. Summary of the invention
[0008] In order to solve the above problems, the present invention proposes a target tracking method, device and medium based on asymmetric background perception.
[0009] The specific plan is as follows:
[0010] A target tracking method based on asymmetric background perception comprises the following steps:
[0011] S1: receiving the first frame training sample and the first frame target template corresponding to the input target to be tracked, and adding the first frame target template to the historical multimodal target pool;
[0012] S2: Based on the training sample of the first frame, the asymmetric background perception matrix of the first frame is calculated; based on the asymmetric background perception matrix of the first frame, the filter introducing the spatial mask constraint of the first frame is calculated, and a filter group containing multiple channels of the first frame is constructed;
[0013] S3: When the t-th frame sample is received, its multi-channel features are extracted;
[0014] S4: Calculate the correlation filter response of the t-th frame sample according to the multi-channel features of the t-th frame sample and the filter group of the t-1-th frame, and obtain the optimal position of the target in the t-th frame sample with the maximum value of the correlation filter response;
[0015] S5: Calculate the optimal scale of the target in the sample of the t-th frame, and determine the optimal target size and the optimal sample size based on the optimal scale;
[0016] S6: Determine the best candidate target for the tth frame according to the best target position and the best target size;
[0017] S7: Calculate the similarity between the best candidate target of the t-th frame and each target in the historical multimodal target pool. If the maximum similarity is less than the similarity threshold, set the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool not to be updated; otherwise, update the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool.
[0018] Furthermore, the calculation formula for the value of the matrix element m in the asymmetric background perception matrix is:
[0019]
[0020]
[0021]
[0022] in, Represents pixel z p is the probability of a pixel in the target, represents the test sample of the tth frame; p(z p |m=1) represents the prior probability of target space movement, p (t-1) represents the position of the target in the t-1th frame, σ represents the standard deviation of the Gaussian window, and |.|2 represents the L2 norm; Represents the color likelihood probability, which is represented by the color histogram c = {c o ,c b} Back-projected to the spatial pixel point to obtain, c o represents the target color histogram, c b Represents the background color histogram, and α represents the limit value of the matrix elements.
[0023] Furthermore, the color likelihood probability The solution is:
[0024]
[0025] in, Represents a one-hot encoded vector, representing pixel z p The extracted color feature vector, which is at position k(z p ) is 1, and the values of other positions are 0, k(z p ) represents pixel z p Mapped to the number of the corresponding color feature in the color histogram; β represents the regression filter of the color histogram.
[0026] Furthermore, the solution method for the regression filter β of the color histogram is:
[0027] Linear regression is performed on the color features of each pixel in the target and background areas, and the objective function is:
[0028]
[0029] Among them, λ represents the ridge regression balance parameter, ||.||2 represents the L2 norm;
[0030] The solution of the above formula is: Where j represents the color number, N j Indicates the total number of colors.
[0031] Furthermore, the calculation formula of the filter that introduces the spatial mask constraint is:
[0032]
[0033] Among them, G d represents the filter that introduces spatial mask constraints, Represents G d The frequency domain signal, F d represents multi-channel features, Indicates F d The frequency domain signal, Indicates F d The frequency domain signal of the conjugate matrix of , ⊙ represents the point multiplication operator, Y represents the designed target label, represents the frequency domain signal of the conjugate matrix of Y, β represents the coefficient of the quadratic penalty function, mat represents the operator that converts the vector into a matrix, Indicates size The discrete Fourier transform matrix, D represents the number of pixels, represents the Kronecker product operator, P m represents a D×D diagonal matrix, h d Indicates H d The vector form of H d represents the filter without spatial mask constraint, L d represents the Lagrange multiplier, Indicates L d The frequency domain signal, d represents the channel number, N d represents the total number of channels, the division sign represents the element-by-element point-by-point division operation, and i represents the number of iterations.
[0034] Furthermore, the process of obtaining the optimal scale of the target in step S5 includes:
[0035] Constructing the objective function of the adaptive filter for:
[0036]
[0037] Among them, y s represents the scale training label, σ2 represents the standard deviation of the expected response Gaussian function, n represents the scale number, N represents the total number of scales, and s k represents the kth channel scale filter, represents the kth channel scale filter s k The reflected signal, λ represents the weight coefficient, k represents the channel number, K represents the total number of channels, ||.||2 represents the L2 norm, Represents the scale feature of the kth channel of the target in the t-1th frame;
[0038] Rewrite the above formula into frequency domain expression, namely:
[0039]
[0040] in, Indicates k The frequency domain signal, Indicates k The frequency domain signal of the conjugate matrix of ; express The frequency domain signal of Represents y s The frequency domain signal of
[0041] make The scale filter is obtained as:
[0042]
[0043] in, express The frequency domain signal of the conjugate matrix of ;
[0044] The scale response of the sample obtained based on the scale filter is:
[0045]
[0046] Among them, ifft represents the one-dimensional inverse Fourier transform operator, real represents the real part operator, Represents the scale feature of the kth channel of the tth frame sample;
[0047] The scale corresponding to the position with the largest scale response is selected as the optimal scale of the target.
[0048] Furthermore, in step S7, the updating method of the t-th frame training sample, the t-th frame asymmetric background perception matrix and the t-th frame filter bank is:
[0049] Re-acquire the t-th frame sample with the best candidate target as the center and the best sample size as the size, and obtain the t-th frame training sample based on the t-th frame sample and the t-1-th frame training sample;
[0050] Counting the foreground histogram and background histogram of the t-th frame training sample, and combining the foreground histogram and background histogram of the t-1-th frame training sample, to obtain the foreground histogram and background histogram of the t-th frame training sample;
[0051] Based on the foreground histogram and background histogram of the training sample of the t-th frame, update the asymmetric background perception matrix of the t-th frame;
[0052] Based on the asymmetric background perception matrix of the t-th frame, the filter of the t-th frame is updated, and the updated filter is used to form a filter group of the t-th frame.
[0053] Furthermore, the updating method of the historical multimodal target pool in step S7 is: replacing the target template with the smallest similarity to the best candidate target in the historical multimodal target pool with the best candidate target in the tth frame.
[0054] A target tracking terminal device based on asymmetric background perception includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method described above in the embodiment of the present invention are implemented.
[0055] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above in an embodiment of the present invention are implemented.
[0056] The present invention adopts the above technical solution to solve the problems existing in the correlation filtering in the prior art and improve the accuracy of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Shown is a flow chart of Embodiment 1 of the present invention.
[0058] Figure 2 Shown is a schematic diagram of background perception and asymmetric background perception in this embodiment.
[0059] Figure 3 Shown is a schematic diagram of the model structure in this embodiment.
[0060] Figure 4 FIG. 4 is a schematic diagram of the scale adaptive filtering process in this embodiment. DETAILED DESCRIPTION
[0061] To further illustrate various embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention, which are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, ordinary technicians in this field should be able to understand other possible implementations and advantages of the present invention.
[0062] The present invention will now be further described with reference to the accompanying drawings and specific implementation methods.
[0063] Embodiment 1:
[0064] The embodiment of the present invention provides a target tracking method based on asymmetric background perception, such as Figure 1 As shown, the method comprises the following steps:
[0065] S1: receiving the first frame training sample and the first frame target template corresponding to the input target to be tracked, and adding the first frame target template to the historical multimodal target pool;
[0066] S2: Based on the training sample of the first frame, the asymmetric background perception matrix of the first frame is calculated; based on the asymmetric background perception matrix of the first frame, the filter introducing the spatial mask constraint of the first frame is calculated, and a filter group containing multiple channels of the first frame is constructed;
[0067] S3: When the t-th frame sample is received, its multi-channel features are extracted;
[0068] S4: Calculate the correlation filter response of the t-th frame sample according to the multi-channel features of the t-th frame sample and the filter group of the t-1-th frame, and obtain the optimal position of the target in the t-th frame sample with the maximum value of the correlation filter response;
[0069] S5: Calculate the optimal scale of the target in the sample of the t-th frame, and determine the optimal target size and the optimal sample size based on the optimal scale;
[0070] S6: Determine the best candidate target for the tth frame according to the best target position and the best target size;
[0071] S7: Calculate the similarity between the best candidate target of the t-th frame and each target in the historical multimodal target pool. If the maximum similarity is less than the similarity threshold, set the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool not to be updated; otherwise, update the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool.
[0072] The key technologies involved in the method are described below.
[0073] (I) Asymmetric background perception matrix
[0074] The asymmetric background perception matrix based on color statistical features designs an irregular mask matrix according to the color histogram of the target to assign different attention weights to the space, so as to express the possibility that the spatial pixel is the pixel contained in the target.
[0075] Assume that the current frame test sample is Pixel The probability of being a pixel in the target is:
[0076]
[0077] Where m = 1 means that The pixels come from the target, m = 0 means z p From the background. The symbol ∝ indicates a proportional relationship; represents the spatial prior probability of the target moving in space (p (t-1) represents the target position of the previous frame; σ is the standard deviation of the Gaussian window). Represents the color likelihood probability, which is represented by the color histogram c = {c o ,c b}Reverse projection to the spatial pixel point to obtain, is the target color histogram, is the background color histogram. The specific solution method is as follows:
[0078]
[0079] in is a one-hot encoded vector (the vector at position k(z p ) is 1, and the values of other positions are 0, k(z p ) represents pixel z p Mapped to the number of the corresponding color feature in the color histogram), indicating pixel z p The extracted color feature vector; Regression filter representing the color histogram is solved as follows:
[0080] On Target and background area Linear regression is performed on the color features of each pixel, and its objective function is:
[0081]
[0082] Where λ is the ridge regression equilibrium parameter, is a vector The L2 norm of .
[0083] The solution of the above formula is:
[0084]
[0085] Among them, j represents the color number, N j Indicates the total number of colors.
[0086] The spatial attention is obtained by binarizing the probability in formula (1) The matrix element m∈{0,1} in the asymmetric background perception matrix is defined as:
[0087]
[0088] Where α represents the set limit value of the matrix element.
[0089] Figure 2 Schematic diagram of traditional background perception and asymmetric background perception proposed in this embodiment. Figure 2 It can be seen that constructing an irregular sampling mechanism based on target shape information can more effectively increase the difference between the background and the target, and is more adaptable to changes in the appearance shape of the target.
[0090] (II) Filter
[0091] This embodiment proposes a model such as Figure 3 As shown in FIG. 1 , an asymmetric mask matrix M is first introduced to construct an asymmetric background-aware correlation filtering framework. On this basis, two sets of filters are designed, one for target tracking and positioning, and the other for target size estimation. Figure 3 The upper branch is the target positioning branch, which introduces boundary suppression regularization terms, spatiotemporal outlier suppression regularization terms, and intra-class distractor suppression regularization terms to construct a spatiotemporal regularized correlation filter. The boundary suppression regularization term uses a fixed inverse Gaussian shape enhancement coefficient in SRDCF to suppress the boundary effect caused by the periodic boundary condition assumption; the spatiotemporal outlier suppression regularization term is used to suppress the pixel points with drastic changes in the corresponding changes between the previous and next frames, and improve the tracking drift caused by the out-of-plane rotation and drastic deformation of the target; the intra-class distractor suppression regularization term determines the interference source in the video through the secondary peak detection in the previous frame response, and designs the distractor suppression regularization term based on the interference source position. Figure 3 The lower branch is the scale estimation branch. In this stage, this embodiment samples samples at multiple scales, extracts features, and pulls the features into column vectors. On this basis, a one-dimensional filter is designed to determine the optimal scale of the target. The technical details of scale estimation will be elaborated in detail later.
[0092] DCF-based methods often rely on manual features for image description. However, manual features are not able to distinguish scenes with in-plane rotation and cluttered backgrounds. To this end, this embodiment introduces the imagenet-vgg-2048 network and uses the MatConvNet library to extract deep features. On this basis, an asymmetric background perception matrix M is introduced to construct an asymmetric background perception correlation filtering framework. For each channel filter H d Introducing spatial attention constraints, namely G d =M⊙H d , the tracking model objective function proposed in this embodiment is:
[0093]
[0094] in Representation Matrix The L2 norm of ; It is the multi-channel feature extracted from the training sample; Indicates the target label of the design; H d represents a filter without spatial mask constraints; G d represents a filter that introduces a spatial mask constraint; represents the spatial domain related operator; ⊙ represents the point product operator; st represents the constraint; represents the spatiotemporal regularization factor, which is defined as:
[0095]
[0096] Where S b represents the boundary effect suppression factor, which is an inverse Gaussian function of a fixed shape and is used to suppress the boundary effect; R stands for abnormal suppressor factor, (t-1) [Δ t-1 ] indicates that it will respond to R (t-1) The maximum value of the shift operator [Δ t-1 ] is the response distribution after moving to the center of the search space, R (t) [Δ t ] indicates that it will respond to R (t) The maximum value of the circular shift operator [Δ t ] is the response distribution after moving to the center of the search space, where the shift distance Δ t By response R (t) The relative distance between the maximum value position and the center position is determined; S i =I i [Δ t ] represents the interference suppression factor, where I i represents the interference detection matrix, whose elements are (where V is the set of pixels near the interferer), Ii [Δ t ] indicates that I s By cyclic shift operator [Δ t ]The matrix after moving to the center of the search space.
[0097] 1. Filter solution
[0098] Introducing the Lagrange multiplier L d and quadratic penalty constraint, the augmented Lagrangian function of the objective function is:
[0099]
[0100] Where tr(X) means finding the trace of matrix X; β is the coefficient of the quadratic penalty function. By G d After reverse arrangement row by row, circular shift 1 bit, and then reverse arrangement column by column, circular shift 1 bit, for example, if but
[0101] Using the spatial convolution theorem, equation (8) is transformed into:
[0102]
[0103] in represents the frequency domain signal of X, fft2 represents the two-dimensional Fourier transform operator, and ⊙ represents the point multiplication operator.
[0104] Rewrite equation (9) into vector form:
[0105]
[0106] in h d for H d The vector form of (in is the Kronecker product operator, Indicates size The discrete Fourier transform matrix, D = h × w), is a diagonal matrix, and its diagonal elements are Represents the vector elements Operator that fills the diagonal elements of a matrix with zeros.
[0107] Each subproblem can be solved by iterative minimization using the alternating direction method:
[0108]
[0109] in and They represent the objective function of each sub-problem respectively.
[0110] (1) Subproblem solving
[0111] The augmented Lagrangian function is The relevant terms consist of The sub-objective function is:
[0112]
[0113] The objective function with respect to the variable Taking the derivative and setting it to zero, we get:
[0114]
[0115] against Combining like terms, we get:
[0116]
[0117] The division sign in the above formula represents the element-by-element point-by-point division operation.
[0118] Transforming equation (14) into matrix form, we have:
[0119]
[0120] Where mat represents the operator that converts the vector into a matrix, the superscript ︿ represents the frequency domain signal, and the superscript * in the upper right corner represents the conjugate of the matrix, that is, represents the frequency domain signal of the conjugate matrix of matrix Y, Denotes the matrix F d The frequency domain signal of the conjugate matrix is Denotes the matrix F d frequency domain signal.
[0121] (2)h d Subproblem solving
[0122] The augmented Lagrangian function is d The relevant terms form about h d The sub-objective function is:
[0123]
[0124] The objective function is about the variable h d Taking the derivative and setting it to zero, we get:
[0125]
[0126] Where F satisfies F HF=DI D , is the identity matrix.
[0127] Then we have:
[0128]
[0129] Assumptions Then F satisfies F H x=Dvec(ifft2(X)) (ifft2 represents the two-dimensional inverse Fourier transform operator), then:
[0130]
[0131] have to:
[0132]
[0133] (3) Subproblem solving
[0134] The objective function of the subproblem is:
[0135]
[0136] Using the gradient ascent method we get:
[0137]
[0138] 2. Correlation filter response
[0139] This embodiment adds a weight based on channel attention in the calculation of the correlation filter response. The calculation process of the correlation filter response of the t-th frame sample is described below:
[0140] (1) Calculate the weight of the channel learning stage of the t-th frame based on the t-1th frame training sample
[0141]
[0142] (2) Calculate the channel-related filter response of the t-th frame sample And determine the secondary peak ρ based on the t-th frame response graph max2 With the main peak ρ max1 , according to the secondary peak ρ max2 and the main peak ρ max1 Calculate the weight of the channel detection stage of the tth frame
[0143] (3) Based on the weight of the t-th frame channel learning stage and the weight of the channel detection stage of the tth frame Calculate the channel weight of the tth frame
[0144] (4) Based on the channel weight of the tth frame and the channel attention weight of the t-1th frame Calculate the channel attention weight of the tth frame Linear interpolation learning rate η and first frame channel attention weight Need to be set up in advance;
[0145] (5) Combined with the channel attention weight of the tth frame Calculate the relevant filter response of the t-th frame sample based on the multi-channel features of the t-th frame sample and the filter bank of the t-1-th frame
[0146] 3. Best scale estimation
[0147] This embodiment designs a scale adaptive filter to effectively perceive the scale change of the target during motion. Assume that the target size of the previous frame is First, the target samples are sampled at different scales, with a scale range of At each scale, the target size is And obtain image features (which can be grayscale features, HOG features and other multi-channel visual features) at different scales (assuming there are N scales in total). Then pull the image features at the nth scale into a column vector and weight the scale prior function (the scale prior function is a Gaussian function, σ1 represents the standard deviation of the scale prior function), forming a multi-scale feature matrix Divide the matrix into blocks by rows, and we get in Represents the scale feature from the target sample, such as Figure 4 shown.
[0148] The objective function of the scale adaptive filter is designed as:
[0149]
[0150] in Represents the scale training label, whose elements are defined as σ2 represents the standard deviation of the expected response Gaussian function. represents the kth channel scale filter s k The reflection signal.
[0151] Rewrite equation (23) into a frequency domain expression, namely:
[0152]
[0153] make The scale filter is obtained as:
[0154]
[0155] For new samples, in order to determine their optimal scale, it is also necessary to perform pyramid sampling on the samples, obtain patches of N scales, and then obtain the multi-scale feature matrix Similarly, the matrix is divided into blocks by row, and we get represents the scale feature from the search sample, then its scale response is:
[0156]
[0157] Where ifft represents the one-dimensional inverse Fourier transform operator, and real represents the real part operator.
[0158] Then select the optimal scale s corresponding to the position with the largest scale response b , and then according to the optimal scale s b Determine the optimal target size: And the optimal sample size: 3. Optimal Sample Reliability Evaluation Mechanism
[0159] In the target tracking process, the target template update strategy is crucial. If the template is not updated, the apparent changes of the template cannot be perceived in time. In the case of partial occlusion, motion blur, etc., the overall template is still updated without principle, which will introduce invalid apparent changes. Traditional DCF-based algorithms update the template and filter in every frame, which is prone to drift in the case of strong occlusion.
[0160] In order to effectively solve the above problems, an optimal sample reliability evaluation mechanism is constructed in this embodiment. The mechanism evaluates the reliability of different samples by storing the appearance patches of the target at different times in history, and selects the most reliable sample for tracking. Specifically, when a new frame appears, a filter is first used to obtain the sample with the largest response value, and the HOG feature is extracted from this sample. The extracted HOG features are then compared with the sample features in the multimodal target pool to find the sample with the highest similarity to the hard positive sample at a certain time in history. If its similarity exceeds the set threshold, the corresponding patch is placed in the multimodal target pool and the sample is considered to be reliable. Otherwise, the candidate sample will be considered unreliable, and it should be avoided from being placed in the template pool, and the filter should be stopped from being updated. The following is a brief introduction to the anti-occlusion method using the optimal sample reliability evaluation mechanism.
[0161] First, we construct a multimodal target pool. For the first frame, since there is no historical data, we should fill the multimodal target pool with the target patches of the first frame, that is: t n =vec(P (1))(n=1,2,…,N), where (In order to reduce the amount of calculation in the actual operation process, the first frame target patch After the color to grayscale image transformation, it is converted into matrix form P (1) ) represents the first frame patch, t n represents the nth column vector of the multimodal target pool T. Starting from the second frame, assuming that the optimal sample obtained by the relevant response is (To facilitate subsequent storage and calculation, we will Convert to a grayscale image of the same size as the target patch in the first frame Its column vector b = vec(B)), extract t n and the HOG features of b, as shown in equations (27)-(28):
[0162]
[0163]
[0164] Where HOG represents the Histogram of Directed Gradients extraction operator.
[0165] The following formula can be used to determine whether the target is blocked:
[0166]
[0167] Where τ is a threshold value in the range [0,1]. When the similarity is greater than the preset threshold, it means that the target is not occluded. At this time, b is updated to the template pool, and the second to Nth templates with the lowest similarity to the target in the template pool are eliminated. If the similarity is less than the threshold, the target is considered to be occluded, and the training samples, filters, multimodal templates, and target and background color histograms will not be updated.
[0168] The training sample update is shown in formula (30):
[0169]
[0170] in, is the training sample of the tth frame, is the previous frame training sample, For The best sample of the current frame is obtained by expanding the center of , and η is the learning rate.
[0171] The embodiments of the present invention have the following beneficial effects:
[0172] (1) The contour information of the target is perceived based on the target's color likelihood probability, and an asymmetric sampling framework is proposed that is different from the traditional correlation filter rule sampling framework to further improve the discrimination between background and target samples;
[0173] (2) A deep neural network is introduced into the correlation filtering framework to extract the deep features of the target and adaptively assign channel weights to each feature channel to suppress the negative impact of intra-class interference on tracking;
[0174] (3) A multimodal template pool is constructed to evaluate the optimal candidate samples in each frame, which not only fully exploits the diversity of targets but also solves the tracking drift and failure problems caused by invalid appearance changes in scenes such as occlusion and intense motion.
[0175] Embodiment 2:
[0176] The present invention also provides a target tracking terminal device based on asymmetric background perception, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the above-mentioned method embodiment of embodiment 1 of the present invention when executing the computer program.
[0177] Further, as an executable solution, the target tracking terminal device based on asymmetric background perception can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The target tracking terminal device based on asymmetric background perception may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the composition structure of the target tracking terminal device based on asymmetric background perception is only an example of a target tracking terminal device based on asymmetric background perception, and does not constitute a limitation on the target tracking terminal device based on asymmetric background perception. It may include more or less components than the above, or a combination of certain components, or different components. For example, the target tracking terminal device based on asymmetric background perception may also include input and output devices, network access devices, buses, etc., and the embodiments of the present invention do not limit this.
[0178] Further, as an executable solution, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the target tracking terminal device based on asymmetric background perception, and uses various interfaces and lines to connect various parts of the entire target tracking terminal device based on asymmetric background perception.
[0179] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the target tracking terminal device based on asymmetric background perception by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and an application required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0180] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method in the embodiment of the present invention are implemented.
[0181] If the module / unit integrated in the target tracking terminal device based on asymmetric background perception is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory) and software distribution medium, etc.
[0182] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, it should be understood by those skilled in the art that various changes may be made to the present invention in form and details without departing from the spirit and scope of the present invention as defined by the appended claims, all of which are within the scope of protection of the present invention.
Claims
1. A target tracking method based on asymmetric background perception, characterized in that: The following steps are involved: S1: receiving the first frame training sample and the first frame target template corresponding to the input target to be tracked, and adding the first frame target template to the historical multimodal target pool; S2: Based on the training sample of the first frame, the asymmetric background perception matrix of the first frame is calculated; based on the asymmetric background perception matrix of the first frame, the filter introducing the spatial mask constraint of the first frame is calculated, and a filter group containing multiple channels of the first frame is constructed; The calculation formula for the value of the matrix element m in the asymmetric background perception matrix is: in, Represents pixel z p is the probability of a pixel in the target, represents the test sample of the tth frame; p(z p |m=1) represents the prior probability of target space movement, p (t-1) represents the position of the target in the t-1th frame, σ represents the standard deviation of the Gaussian window, and ||.||2 represents the L2 norm; Represents the color likelihood probability, which is represented by the color histogram c = {c o ,c b }Reverse projection to the spatial pixel point to obtain, c o represents the target color histogram, c b represents the background color histogram, and α represents the limit value of the matrix element; The calculation formula of the filter that introduces the spatial mask constraint is: Among them, G d represents the filter that introduces spatial mask constraints, Represents G d The frequency domain signal, F d represents multi-channel features, Indicates F d The frequency domain signal, Indicates F d The frequency domain signal of the conjugate matrix of , ⊙ represents the point multiplication operator, Y represents the designed target label, represents the frequency domain signal of the conjugate matrix of Y, β represents the coefficient of the quadratic penalty function, mat represents the operator that converts the vector into a matrix, Indicates size The discrete Fourier transform matrix, D represents the number of pixels, represents the Kronecker product operator, P m represents a D×D diagonal matrix, h d Indicates H d The vector form of H d represents the filter without spatial mask constraint, L d represents the Lagrange multiplier, Indicates L d The frequency domain signal, d represents the channel number, N d represents the total number of channels, the division sign represents the element-by-element point-by-point division operation, and i represents the number of iterations; S3: When the t-th frame sample is received, its multi-channel features are extracted; S4: Calculate the correlation filter response of the t-th frame sample according to the multi-channel features of the t-th frame sample and the filter group of the t-1-th frame, and obtain the optimal position of the target in the t-th frame sample with the maximum value of the correlation filter response; S5: Calculate the optimal scale of the target in the sample of the t-th frame, and determine the optimal target size and the optimal sample size based on the optimal scale; S6: Determine the best candidate target for the tth frame according to the best target position and the best target size; S7: Calculate the similarity between the best candidate target of the t-th frame and each target in the historical multimodal target pool. If the maximum similarity is less than the similarity threshold, set the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool not to be updated; otherwise, update the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter group and the historical multimodal target pool.
2. The target tracking method based on asymmetric background perception according to claim 1, characterized in that: Color likelihood probability The solution is: in, Represents a one-hot encoded vector, representing pixel z p The extracted color feature vector, which is at position k(z p ) is 1, and the values of other positions are 0, k(z p ) represents pixel z p Mapped to the number of the corresponding color feature in the color histogram; β represents the regression filter of the color histogram.
3. The target tracking method based on asymmetric background perception according to claim 2, characterized in that: The solution method for the regression filter β of the color histogram is: Linear regression is performed on the color features of each pixel in the target and background areas, and the objective function is: Among them, λ represents the ridge regression balance parameter, ||.||2 represents the L2 norm; The solution of the above formula is: Where j represents the color number, N j Indicates the total number of colors.
4. The target tracking method based on asymmetric background perception according to claim 1, characterized in that: The calculation process of the correlation filter response of the t-th frame sample is: Calculate the weight of the channel learning stage of the t-th frame based on the t-1-th frame training sample; Calculate the relevant filter response of each channel of the t-th frame sample, determine the secondary peak and the main peak based on the t-th frame response graph, and calculate the weight of the t-th frame channel detection stage according to the secondary peak and the main peak; Calculate the channel weight of the t-th frame based on the channel weight of the t-th frame learning phase and the channel weight of the t-th frame detection phase; Based on the channel weight of the t-th frame and the channel attention weight of the t-1-th frame, calculate the channel attention weight of the t-th frame; Combined with the channel attention weights of the t-th frame, the relevant filter response of the t-th frame sample is calculated according to the multi-channel features of the t-th frame sample and the filter group of the t-1th frame.
5. The target tracking method based on asymmetric background perception according to claim 1, characterized in that: The process of obtaining the optimal scale of the target in step S5 includes: Constructing the objective function of the adaptive filter for: Among them, y s represents the scale training label, σ2 represents the standard deviation of the expected response Gaussian function, n represents the scale number, N represents the total number of scales, and s k represents the kth channel scale filter, represents the kth channel scale filter s k The reflected signal, λ represents the weight coefficient, k represents the channel number, K represents the total number of channels, ||.||2 represents the L2 norm, and f k x Represents the scale feature of the kth channel of the target in the t-1th frame; Rewrite the above formula into frequency domain expression, namely: in, Indicates k The frequency domain signal, Indicates k The frequency domain signal of the conjugate matrix of ; represents f k x The frequency domain signal of Represents y s The frequency domain signal of make The scale filter is obtained as: in, represents f k x The frequency domain signal of the conjugate matrix of ; The scale response of the sample obtained based on the scale filter is: Among them, ifft represents the one-dimensional inverse Fourier transform operator, real represents the real part operator, Represents the scale feature of the kth channel of the tth frame sample; The scale corresponding to the position with the largest scale response is selected as the optimal scale of the target.
6. The target tracking method based on asymmetric background perception according to claim 1, characterized in that: In step S7, the updating method of the t-th frame training sample, the t-th frame asymmetric background perception matrix, the t-th frame filter bank and the historical multimodal target pool is: Re-acquire the t-th frame sample with the best candidate target as the center and the best sample size as the size, and obtain the t-th frame training sample based on the t-th frame sample and the t-1-th frame training sample; Counting the foreground histogram and background histogram of the t-th frame training sample, and combining the foreground histogram and background histogram of the t-1-th frame training sample, to obtain the foreground histogram and background histogram of the t-th frame training sample; Based on the foreground histogram and background histogram of the training sample of the t-th frame, update the asymmetric background perception matrix of the t-th frame; Based on the asymmetric background perception matrix of the t-th frame, the filter of the t-th frame is updated, and the updated filter is used to form a filter group of the t-th frame; The best candidate target in the tth frame is used to replace the target template with the smallest similarity to the best candidate target in the historical multimodal target pool.
7. A target tracking terminal device based on asymmetric background perception, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.