A single object tracking method based on accurate bounding box prediction
By adopting a method based on precise bounding box prediction in single-target tracking, combining pixel cross-correlation and channel attention mechanisms, the shortcomings of the twin network in dealing with target similar interference and background noise are solved, and the tracking accuracy and robustness are significantly improved.
Patent Information
- Application Number
- CN202310515531.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2043-05-09
AI Technical Summary
The existing twin network structures are poor in dealing with target similar interference and background noise, making it difficult to maintain the natural spatial structure in the feature map, resulting in insufficient robustness of target tracking when scale changes, rotations and fast motion.
A single-objective tracking method based on precise bounding box prediction is adopted, and the discriminant learning of target-specific features is used through a fractional fusion strategy, combined with pixel cross-correlation and channel attention mechanism, spatial information in the features is extracted and maintained, and the natural spatial structure is maintained through a key point-based bounding box prediction network.
It significantly improves the accuracy and robustness of target tracking, effectively suppresses background noise, maintains the natural spatial structure in the feature map, and ensures accurate tracking when target scale changes, rotates and fast motion.
Smart Images

Figure CN116543019B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, relates to single-object tracking technology, and specifically relates to a single-object tracking method based on accurate bounding box prediction. Background Art
[0002] With the continuous development of science and technology, the degree of social informatization and intelligence is increasing day by day. Mankind has entered the era of big data and informatization, which has brought great convenience to people's lives and also made the research of computer vision more active.
[0003] Visual object tracking is a very important and challenging branch in the field of computer vision. It refers to detecting, extracting, recognizing, and tracking moving objects in an image sequence, obtaining motion parameters such as the position, motion trajectory, speed, and acceleration of the moving object, and processing and analyzing this data to achieve an understanding of the behavior of the moving object and complete high-level video analysis tasks. It is widely used in modern military, video surveillance, autonomous driving, medical diagnosis, and other fields, and has important research value and practical significance.
[0004] Although object tracking technology has been applied in many fields, in the actual tracking process, due to uncontrollable factors, it still faces many challenges, such as the object having illumination changes, motion blur, rotation, interference from similar objects, low resolution, occlusion, shape changes, illumination changes, etc. Therefore, in order to solve the difficulties and challenges encountered in the object tracking process and be better applied in multiple fields, researching and designing a high-precision and real-time object tracking algorithm has important value and far-reaching influence.
[0005] In recent years, with the continuous development and application of deep learning technology, object tracking algorithms based on discriminant models have also been evolving, from object tracking algorithms based on correlation filtering to object tracking algorithms based on deep learning, continuously improving the accuracy, real-time performance, and robustness of the tracking algorithm. The tracker based on the Siamese network has received extensive attention from researchers due to its fast speed and high accuracy. Summary of the Invention
[0006] The object of the present invention is to propose a single-object tracking method based on accurate bounding box prediction in view of the current Siamese network structure lacking background features of specific targets, being unable to effectively identify target analogue interference and reduce the influence of background noise. This method uses discriminative learning of target-specific features through a score fusion strategy to help the Siamese network better handle interference and noise, and effectively extracts and maintains spatial information in features through a strategy of fusing pixel cross-correlation and channel attention mechanism; through a key-point style bounding box prediction network, it can effectively maintain the natural spatial structure in the feature map, and avoid encoding spatial information into channels, improving the robustness when the target undergoes scale changes, rotations and rapid movements. Specifically, it includes the following steps:
[0007] (1) Construct a network model based on accurate bounding box prediction and perform offline training on this model;
[0008] (1a) Input a video sequence, and select two random template frames F ref and test frame F test with an interval of less than 50 frames;
[0009] (1b) Crop the template frame F ref into an image twice the size of the given annotated bounding box as the input of the template branch, and perform translation, flipping, scaling, color change and blurring processing on the image obtained by cropping the test frame F test centered on the annotated bounding box as the input of the search branch, and calculate through the following formula
[0010]
[0011]
[0012]
[0013] A region with [c x ,c y as the center and size [h, w] can be obtained, where respectively represent the abscissa value and ordinate value of the center point of the given annotated bounding box and the length and width of the annotated bounding box, and are two scalar factors, representing scale and center respectively, and N and U represent two-dimensional standard normal distribution random variables and two-dimensional uniform random variables respectively;
[0014] (1c) Convert the prediction output result of the target bounding box into coordinates in the format of leftmost, topmost, rightmost, and bottommost, and compare it with the coordinate values of the given annotated bounding box to obtain the total loss:
[0015] L = L box + λLmask
[0016] Among them, L box represents the mean squared error, and L mask represents the cross-entropy loss, and λ represents the weight coefficient;
[0017] (2) Load the network model of the initial tracking algorithm and initialize the network model of the accurate bounding box prediction algorithm for offline training;
[0018] (3) Optimize the coordinates of the predicted bounding box, perform pixel cross-correlation operations on the features of the extracted search image and template image, and perform squeezing and activation operations on the features after pixel cross-correlation through the channel attention mechanism to obtain response features. The specific steps are as follows:
[0019] (3a) Input the template image features of and 0 the search image features of 0 where C represents the number of feature channels, and H 0 ×W 0 decompose the template image feature K into H ×W smaller convolutional kernels
[0020]
[0021]
[0022] (3b) Generate channel-based statistical information through global average pooling operations, and compress the global spatial information into the channel descriptor. The statistic z ∈ R C is obtained by performing an F c (.) contraction operation on the spatial dimensions H × W of the feature map u sq . Then, the c-th element of z is calculated as
[0023]
[0024] where i represents the i-th row of the feature map u c and j represents the j-th column of the feature map u c ;
[0025] (3c) Generate weights s for each feature channel through the parameter w. The whole process can be described as
[0026] s = F ex (z, w) = σ(w 2δ(w 1 z))
[0027]
[0028] δ(x) = max(0, x)
[0029] where F ex (.) represents an extraction operation, σ(x) represents the Sigmoid activation function, δ(x) represents the ReLU activation function, z represents the compressed feature information, respectively represent the first and second layers of the fully connected layer, where L represents the number of channels of the feature and r represents the feature compression ratio factor;
[0030] (3d) By multiplying each learned channel attention weight s c with the input feature u of the backbone c to obtain the output feature is
[0031]
[0032] where F sc (u c , s c ) represents the channel multiplication between the attention weight s c and the feature map ;
[0033] (4) Calculate the heatmap information of the upper left and lower right points of the target in the response feature, convert it through the probability density function to obtain the predicted bounding box of the target, and update the predicted result of the target bounding box in the initial tracking algorithm to complete the positioning and tracking of the target in the entire video sequence. The specific calculation method is
[0034]
[0035] where h n,m represents the element corresponding to the m-th column and n-th row in the normalized heatmap of size W h ×H h , m represents the m-th column of the heatmap, n represents the n-th row of the heatmap, and p = (p x , p y ) represents the position of the upper left or lower right point of the target.
[0036] The innovation of the present invention is to propose a more flexible, accurate and computationally efficient bounding box prediction module; by fusing the pixel cross-correlation and channel attention mechanisms, effectively extract and maintain the spatial information in the features; adopt a key-point-based bounding box prediction network to effectively suppress background noise and maintain the natural spatial structure in the feature map, significantly improving the bounding box prediction quality of the tracker.
[0037] Advantages of the present invention: effectively solve the problem of target drift during target tracking when the target undergoes appearance changes, rotation, and motion blur; improve the robustness to scale changes and rotation of the target; significantly improve the tracking accuracy on the premise of ensuring real-time tracking speed.
[0038] The present invention is mainly verified by means of simulation experiments, and all steps and conclusions are verified correctly on the open-source target tracking algorithm framework based on pytracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the flowchart of the present invention;
[0040] Figure 2 is the network framework diagram of the present invention;
[0041] Figure 3 is the pixel cross-correlation schematic diagram of the present invention;
[0042] Figure 4 is the structural diagram of the key-point type bounding box prediction network of the present invention;
[0043] Figure 5 is the tracking results of some sequences on the OTB100 dataset using different methods. DETAILED DESCRIPTION OF THE INVENTION
[0044] Referring to Figure 1 , the present invention is a single-object tracking method based on accurate bounding box prediction, and the specific steps are as follows:
[0045] (1) Construct a network model based on accurate bounding box prediction and perform offline training on this model;
[0046] (1a) Input a video sequence, and select two random template frames F ref and test frame F test with an interval of less than 50 frames;
[0047] (1b) Crop the template frame F ref into an image twice the size of the given annotation bounding box as the input of the template branch, and perform translation, flipping, scaling, color change, and blurring processing on the image obtained by cropping the test frame F test centered on the annotation bounding box as the input of the search branch, and calculate through the following formula
[0048]
[0049]
[0050]
[0051] We can obtain a region centered at [c x ,c y with size [h, w], where respectively represent the abscissa value and ordinate value of the center point of the given annotated bounding box and the length and width of the annotated bounding box, and are two scalar factors, representing scale and center respectively, N and U represent two-dimensional standard normal distribution random variables and two-dimensional uniform random variables;
[0052] (1c) Convert the predicted output result of the target bounding box into coordinates in the leftmost, topmost, rightmost, and bottommost format, and compare it with the coordinate values of the given annotated bounding box to obtain the total loss:
[0053] L = L box + λL mask
[0054] where, L box represents the mean squared error, L mask represents the cross-entropy loss, and λ represents the weight coefficient;
[0055] (2) Load the network model of the initial tracking algorithm and initialize the network model of the algorithm based on accurate bounding box prediction for offline training;
[0056] (3) Optimize the coordinates of the predicted bounding box, perform pixel cross-correlation operations on the features of the extracted search image and template image, and perform squeezing and activation operations on the features after pixel cross-correlation through the channel attention mechanism to obtain response features. The specific steps are as follows:
[0057] (3a) Input template image features and search image features, where C represents the number of feature channels, H 0 , W 0 respectively represent the length and width of the template image features, H and W respectively represent the length and width of the search image features. Decompose the template image feature K into H 0 × W 0 smaller convolutional kernels and perform correlation calculation with the search image features to obtain the pixel correlation map The whole process can be described as
[0058]
[0059] where, * represents naive cross-correlation, and the subscript j represents the j-th channel;
[0060] (3b) Generate channel-based statistics through global average pooling operation and compress the global spatial information into the channel descriptor. The statistic z ∈ R C By performing an F c contraction operation on the spatial dimensions H×W of the feature map u sq , the c-th element of z is calculated as
[0061]
[0062] where i represents the i-th row of the feature map u c and j represents the j-th column of the feature map u c ;
[0063] (3c) Generate a weight s for each feature channel through the parameter w. The whole process can be described as
[0064] s = F ex (z, w) = σ(w 2 δ(w 1 z))
[0065]
[0066] δ(x) = max(0, x)
[0067] where F ex (.) represents the extraction operation, σ(x) represents the Sigmoid activation function, δ(x) represents the ReLU activation function, z represents the contracted feature information, represent the first and second layers of the fully connected layer respectively, where L represents the number of channels of the feature and r represents the feature compression ratio factor;
[0068] (3d) Multiply the learned attention weight s c for each channel with the input feature u c of the backbone to obtain the output feature which is
[0069]
[0070] where F sc (u c , s c ) represents the channel multiplication between the scalar s c and the feature map ;
[0071] (4) Calculate the heatmap information of the top-left and bottom-right points of the target in the response feature, convert it through the probability density function to obtain the predicted bounding box of the target, and update the predicted result of the target bounding box in the initial tracking algorithm to complete the localization and tracking of the target in the entire video sequence. The specific calculation method is
[0072]
[0073] Among them, h n,m represents the element corresponding to the m-th column and the n-th row in the normalized heatmap of size W h ×H h , where m represents the m-th column of the heatmap, n represents the n-th row of the heatmap, and p = (p x , p y ) represents the position of the upper left or lower right point of the target.
[0074] The effects of the present invention can be further illustrated by the following simulation experiments:
[0075] I. Experimental conditions and content
[0076] Experimental conditions: Some video sequences in the OTB100 dataset are used in the experiment, as Figure 5 shown; the success rate curve and the precision curve are used as evaluation indicators for the experimental results to objectively evaluate the reconstruction results. The success rate curve is plotted according to the area overlap ratio IoU (Intersection over Union) between the bounding box obtained by the tracking algorithm and the accurate bounding box manually annotated, and its calculation formula is:
[0077]
[0078] where Box P is the target bounding box predicted by the tracking algorithm, and Box G is the true bounding box of the target. A threshold T is set. When the success rate of a certain frame is greater than T, the tracking of this frame is considered successful. The success rate curve reflects the percentage of video frames with a bounding box overlap ratio greater than the given threshold, and can better describe the closeness between the target scale predicted by the tracking algorithm and the true scale. The precision curve is plotted according to the Euclidean distance error between the center of the target bounding box obtained by the tracking algorithm and the center of the accurate bounding box manually annotated, and its calculation formula is:
[0079]
[0080] where (x P , y P ) is the center position of the target bounding box predicted by the tracking algorithm, and (x G , y G ) is the center position of the accurate bounding box manually annotated. A threshold is set. Only when d < T is the tracking of this frame considered successful. Usually, the value corresponding to 20 pixel points is used as the precision evaluation indicator.
[0081] Experimental content: Under the above conditions, the SiamBAN method and the SiamBAN++ method, which are currently at the leading level in the field of single-object tracking, are compared with the method of the present invention. The tracking comparison results are as Figure 5 shown.
[0082] As can be seen from Figure 5 (a), in the Board sequence, the target moves and rotates rapidly, resulting in motion blur. The SiamBAN method loses track of the target. Only the SiamBAN++ method and the method of the present invention make correct predictions. However, since the SiamBAN++ method uses an RPN-style bounding box prediction network and fails to fully utilize the information contained in the spatial distribution of the feature map, the bounding box prediction is inaccurate. Only the method of the present invention most accurately predicts the position of the target.
[0083] As can be seen from Figure 5 (b), in the Clifbar sequence, only the bounding box predicted by the method of the present invention coincides with the correct bounding box manually annotated. The bounding boxes predicted by the SiamBAN method and the SiamBAN++ method have large differences from the correct bounding box manually annotated.
[0084] As can be seen from Figure 5 (c), in the Ironman sequence, there are strong illumination changes around the target, accompanied by similar object interference and occlusion. Both the SiamBAN method and the SiamBAN++ method show the phenomenon of target drift. Only the method of the present invention makes accurate predictions and successfully tracks the target.
[0085] As can be seen from Figure 5 (d), in the Walking2 sequence, there is similar object interference around the target. Both the SiamBAN method and the SiamBAN++ method lose track of the target. Only the method of the present invention can successfully track the target.
[0086] Table 1 Success rate indicators of different tracking methods for some video sequences under the OTB100 dataset
[0087] Video sequence SiamBAN method SiamBAN++ method Method of the present invention Board 0.474 0.730 0.766 Clifbar 0.473 0.509 0.722 Ironman 0.565 0.520 0.645 Walking2 0.279 0.271 0.347
[0088] Table 1 shows the success rate indicators of each tracking method. The larger the success rate value, the better the tracking effect. It can be seen from the table that the method of the present invention has a significant improvement in the tracking success rate compared with other methods.
[0089] Table 2 Precision indicators of different tracking methods for some video sequences under the OTB100 dataset
[0090] Video sequence SiamBAN method SiamBAN++ method Method of the present invention Board 0.431 0.646 0.699 Clifbar 0.790 0.835 0.908 Ironman 0.802 0.668 0.818 Walking2 0.381 0.373 0.428
[0091] Table 2 presents the precision index of each tracking method, where a higher precision value indicates that the predicted bounding box is closer to the manually annotated bounding box; it can be seen that the precision value corresponding to the method of the present invention is the highest, and the predicted bounding box is closer to the manually annotated bounding box, and this result is consistent with the tracking effect diagram.
[0092] The above experiments show that the pixel cross-correlation and channel attention mechanism module proposed by the present invention can solve the influence of target background noise. At the same time, the proposed key-point type bounding box prediction network can effectively solve the data inconsistency problem in the RPN network head, and also solve the problem of spatial information collapse in the R-CNN network, and can maintain the natural spatial structure in the feature map to achieve accurate positioning of the target bounding box.
Claims
1. A single target tracking method based on accurate bounding box prediction, comprising the following steps: (1) Build a network model based on accurate bounding box prediction and train the model offline; (2) Load the network model of the initial tracking algorithm and initialize the network model based on the precise bounding box prediction algorithm trained offline; (3) Optimize the coordinates of the predicted bounding box, perform pixel cross-correlation operations on the extracted search image and template image features, and use the channel attention mechanism to squeeze and activate the pixel cross-correlation features to obtain the response features. The specific steps are as follows: (3a) Input The template image features and The search image features are decomposed into H0×W0 smaller convolution kernels. Calculate the correlation with the search image features to get the pixel correlation map The whole process can be described as Where * represents naive cross-correlation, and subscript j represents the jth channel; (3b) Generate channel-based statistics through global average pooling operation and compress the global spatial information into the channel descriptor, statistic z∈R C By c The spatial dimension of F is H×W sq (.) is contracted, then the cth element of z is Among them, i represents the feature map u c The i-th row, j represents the feature map u c The jth column of (3c) The weight s is generated for each feature channel through the parameter w. The whole process can be described as s=F ex (z,w)=σ(w2δ(w1z)) δ(x)=max(0,x) Among them, F ex (.) represents the extraction operation, σ(x) represents the Sigmoid activation function, δ(x) represents the ReLU activation function, and z represents the feature information after contraction. Represent the first and second layers of the fully connected layer, respectively, where L represents the number of feature channels and r represents the feature compression scale factor; (3d) By taking the learned attention weight s for each channel c With the input feature u of the backbone c Multiply to get the output feature for Among them, F sc (u c ,s c ) represents the attention weight s c and feature map Channel multiplication between ; (4) Calculate the heat map information of the upper left and lower right points of the target in the response feature, convert the predicted bounding box of the target through the probability density function, and update the predicted result of the target bounding box in the initial tracking algorithm to complete the positioning and tracking of the target in the entire video sequence. The specific calculation method is: Among them, h n,m Indicates size W h ×H h The element corresponding to the mth column and nth row in the normalized heat map, m represents the mth column of the heat map, n represents the nth row of the heat map, p = (p x ,p y ) indicates the position of the upper left or lower right point of the target.
2. According to the single target tracking method based on accurate bounding box prediction according to claim 1, the main feature of step (1) is that the specific steps of offline training of the model are: (1a) Input a video sequence and select a random template frame F with an interval of less than 50 frames between two frames. ref and test frame F test ; (1b) By ref The image cropped to twice the size of the given annotation bounding box is used as the input of the template branch, and the test frame F test The cropped image centered on the annotated bounding box is translated, flipped, scaled, color-changed, and blurred as the input of the search branch, calculated using the following formula It can be obtained by [c x ,c y ] is the center and the size is [h,w], where Respectively represent the horizontal and vertical coordinate values of the center point of the given annotation bounding box and the length and width of the annotation bounding box. and are two scalar factors, representing the scale and center respectively, N and U represent a two-dimensional standard normal distribution random variable and a two-dimensional uniform random variable respectively; (1c) Convert the predicted output of the target bounding box into coordinates in the format of leftmost, topmost, rightmost, and bottommost, and compare them with the coordinate values of the given annotated bounding box to obtain the total loss L=L box +λL mask in, L box represents the mean square error, L mask represents the cross entropy loss, and λ represents the weight coefficient.
3. According to the single target tracking method based on precise bounding box prediction described in claim 1, the main feature of its step (3) is that the fusion of template frame features and search frame features is completed by adopting pixel point cross-correlation, and the introduction of channel attention mechanism can ensure that each correlation graph can be mapped to the information of a local area of the target, avoiding the phenomenon of feature blurring caused by a large correlation window.
4. According to the single target tracking method based on precise bounding box prediction described in claim 1, the main feature of its step (4) is that the heat map is normalized by the probability density function, which can achieve efficient pixel positioning, so that the discrete heat map can more accurately describe the position information of the upper left and lower right points of the target, and predict continuous values from the discrete heat map, effectively avoiding the data inconsistency problem in the RPN network head, solving the problem of spatial information collapse of the R-CNN network, and being able to maintain the natural spatial structure in the feature map, avoiding encoding spatial information into the channel.