Object Tracking Method and System Based on Voting of Probability Maps of Key Point Depth
By outputting the existence probability graph of five key points and the target box generation and voting steps, the problems of noise interference and historical tracking cumulative errors in the existing methods are solved, and more accurate target tracking is achieved.
Patent Information
- Application Number
- CN202210724437.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-24
AI Technical Summary
The existing deep learning-based target tracking methods are easily disturbed by noise, and the key point positioning deviation affects the target positioning accuracy, and relying on historical tracking results leads to cumulative errors, affecting the accuracy of tracking results.
By outputting the existence probability map of five key points, increasing the characteristic robustness, designing the target box generation and voting steps, generating six target boxes without overlap and weighting fusion, and finally determining the target position.
It improves the accuracy and robustness of target tracking, weakens the impact of errors, and enhances the richness of features and the utilization of information.
Smart Images

Figure CN115393391B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision target tracking, and particularly to a target tracking method and system based on voting of a depth existence probability map of key points. Background Art
[0002] Target tracking is one of the basic research directions in the field of computer vision. Its purpose is to predict the position of a given target in subsequent frames after its precise state in the initial frame is given. Currently, it has been widely used in fields such as video surveillance, human-computer interaction, and autonomous driving. With the gradual improvement of deep learning theory and the increasing progress of computing resources, the target tracking method based on deep learning has developed rapidly and achieved excellent results.
[0003] Most of the mainstream deep learning-based target tracking methods adopt a detection-based tracking method, that is, inputting the template frame and the tracking frame into the neural network at the same time, designing a method to learn the similarity between the template frame and the tracking frame, and outputting the target position in the current frame in the form of a bounding box. According to the output type of the neural network, the tracking methods can be classified into two categories: "top-down" and "bottom-up". The "top-down" method outputs multiple possible target bounding boxes and their corresponding probabilities, and determines the final target bounding box by sorting the probabilities. The "bottom-up" method focuses on identifying the sub-features that make up the target box (such as the upper left point, lower right point, center point, etc. of the target box), and fusing the sub-features to generate the target box. The "top-down" method is affected by the need for task-related prior knowledge and complex parameter adjustment steps ("anchor box" method), and the internal contradiction between the "classification" branch and the "regression" branch ("anchor-free" method), resulting in the impact on the method performance and the generalization ability of the method. The "bottom-up" method obtains the target bounding box by finding the key point features of the target. In the past, this method used a designed network to output the existence probability maps of the upper left point and the lower right point that make up the target box, and regarded the position of the maximum value in the existence probability map as the position of this key point, and then obtained the target box. However, this method is easily interfered by noise, and any deviation in the positioning of a key point will directly affect the target positioning. At the same time, most target tracking methods rely on historical tracking results, and the failure of tracking in one frame will bring cumulative errors to subsequent tracking, seriously affecting the accuracy of the tracking results. Summary of the Invention
[0004] To solve the above problems, the present invention provides a target tracking method and system based on voting of a depth existence probability map of key points. By introducing additional key point information (upper left point, lower right point, lower left point, upper right point, and center point), designing a neural network to output the existence probability maps of five key points of the target (i.e., key point generation), and designing two steps of target box generation and target box voting to determine the final output target box.
[0005] Compared with only outputting the probability maps of the upper-left to lower-right double key points, the method of the present invention can fully obtain target information by outputting the probability maps of five key points, increase the robustness of features, weaken the influence of errors. At the same time, due to the increase in information, the generated target boxes increase. By designing a "voting" step to perform weighted fusion on the filtered target boxes, the target position can be obtained more accurately.
[0006] The present invention provides an object tracking method based on voting of the depth probability maps of key points. The specific technical solution is as follows:
[0007] S1: Obtain a template frame and a tracking frame, input them into a depth network model, generate five key point feature maps and output the key point coordinates.
[0008] S2: Based on the coordinates of each key point, perform pairwise combinations to generate non-overlapping target boxes, and screen the generated target boxes.
[0009] S3: Perform voting calculation according to the filtered target boxes to obtain the final target bounding box.
[0010] Further, in step S1, the specific process is as follows:
[0011] S101: Construct a depth network model based on a Siamese or Transformer depth network structure. Through convolutional stacking and upsampling operations on the head of the network model, adjust the feature size so that the head outputs five key point probability maps.
[0012] S102: Construct five all-zero matrices with the same size as the probability maps, and add two-dimensional Gaussian functions at the key point coordinates of each matrix to construct the ground truth.
[0013] S103: Use a binary classification loss function to calculate the difference between the probability maps and the ground truth, and train the network model that outputs five key point probability maps.
[0014] S104: Based on the trained network model, input the template frame and the tracking frame, output five key point probability maps, select the position of the maximum value in the probability maps as the position of the key point, and output the coordinates of the five key points respectively.
[0015] Further, in step S101, the five key point probability maps respectively represent the existence probabilities of the upper-left point, lower-right point, lower-left point, upper-right point and center point of the target.
[0016] Further, in step S2, after generating the target boxes, it further includes calculating the pairwise intersection-over-union matrix of the target boxes, and obtaining the sum of the pairwise intersection-over-unions of each target box and the remaining target boxes for target box screening.
[0017] Further, in step S2, the specific process is as follows:
[0018] S201: Based on the coordinates P of the five key points output i (x i , y i ), i ∈ (A~E), combine them in pairs to generate six non - overlapping target boxes B m (x m , y m , w m , h m ), m ∈ {AD, BC, AE, BE, CE, DE}, where (x m , y m ) represents the center of the bounding box, and (w m , h m ) represents the width and height of the bounding box;
[0019] S202: Calculate the pairwise intersection - over - union matrix of the six possible target boxes B m , sum the matrix horizontally to obtain the sum of the pairwise intersection - over - union of each target box with the other target boxes, which is used for target box screening;
[0020] S203: Set the screening value l, select the top l target boxes with the sum of intersection - over - union, and vote on the selected target boxes.
[0021] Further, the screening value l is half of the number of target boxes obtained in step S201.
[0022] Further, in step S3, the specific process is as follows:
[0023] S301: Calculate the voting weight of each target box participating in the vote, and the voting weight is related to the attributes of the target box itself;
[0024] S302: Weight the target boxes based on the voting weights of the target boxes to obtain the final target bounding box.
[0025] Further, step S301 is specifically:
[0026] Define the quality of the key point as the value corresponding to the point in the probability map of existence. Then the score of each box is the sum of the scores of the two points constituting the box, and use the quality of the normalized target box as the voting weight of the corresponding target box.
[0027] The present invention also provides a target tracking system based on voting of the depth probability map of key points, including a key point generation module, a target box generation module, and a target box voting module;
[0028] The key point generation module obtains the existence probability map of key points based on the template frame and the tracking frame, and obtains the coordinates of five key points according to the existence probability map of key points;
[0029] The target box generation module is used to combine the coordinates of five key points in pairs to generate six non-overlapping target boxes, calculate the pairwise intersection-over-union matrix of the six non-overlapping target boxes, sum the matrix horizontally to obtain the sum of the pairwise intersection-over-union of each target box with the remaining target boxes, and perform screening based on the sum of the intersection-over-union;
[0030] The target box voting module is used to set the voting weights of the target boxes participating in the voting, and weight the target boxes based on the target box voting weights to obtain the final target bounding box.
[0031] Further, the key point generation module inputs the template frame and the tracking frame into a deep learning network model with five key point existence probability maps output by the head to obtain the existence probability map of key points, constructs the ground truth by adding a two-dimensional Gaussian function at the key point coordinates of a all-zero matrix with the same size as the existence probability map, trains the parameters of the deep network model by reducing the loss between the existence probability map and the ground truth, and when tracking, selects the position of the maximum value of each existence probability map as the position of the key point, and outputs the coordinates of five key points.
[0032] The beneficial effects of the present invention are as follows:
[0033] 1. By adopting a network that outputs multiple key point existence probability maps, a deep network model is constructed based on a Siamese or Transformer deep network structure, so that the head of the model outputs five key point existence probability maps, constructs five all-zero matrices with the same size as the existence probability maps, adds a two-dimensional Gaussian function at the key point coordinates of each matrix respectively to construct the ground truth, and uses a binary classification loss function to calculate the difference between the existence probability map and the ground truth, selects the position of the maximum value in the existence probability map as the position of the key point, and outputs the coordinates of five key points respectively. The target information is fully obtained through the existence probability maps of the five key points, the richness and robustness of the features are increased, and the influence of errors is weakened.
[0034] 2. Six non-overlapping target boxes are generated by combining the output five key points in pairs, and are calculated and screened through the pairwise intersection-over-union matrix to remove relatively weak target boxes, and then the voting weights of each target box participating in the voting are set, and the final result is output through weighted calculation based on the voting weights; the final target position reasonably screens and weights the six non-overlapping target boxes generated by the five key points through two steps of target box generation and target box voting, improving the accuracy of the tracking result. Description of the Drawings
[0035] Figure 1It is a schematic diagram of the overall process of the method of the present invention;
[0036] Figure 2 It is a schematic diagram of the key point generation process of the present invention;
[0037] Figure 3 It is a schematic diagram of the pairwise intersection ratio matrix of the target boxes of the present invention. Detailed implementation manners
[0038] In the following description, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0039] Embodiment 1
[0040] Embodiment 1 of the present invention discloses an object tracking method based on voting of key point depth existence probability maps. As Figure 1 shown, the specific steps are as follows:
[0041] S1: Obtain a template frame and a tracking frame, input them into a depth network model, generate five key points and output the targets of each key point.
[0042] As Figure 2 shown, the specific process is as follows:
[0043] S101: Design a method to learn the similarity between the template frame and the tracking frame. Based on a siamese or Transformer depth network structure, construct a depth network model. Through operations such as convolutional stacking and upsampling on the head of the network model, adjust the feature size. Input the template frame and the tracking frame into the neural network simultaneously, so that the head outputs five key point existence probability maps; the siamese depth network structure and the Transformer depth network structure are both existing deep learning network model framework structures.
[0044] The five key point existence probability maps respectively represent the existence probabilities of the upper left point, the lower right point, the lower left point, the upper right point and the center point of the target;
[0045] S102: Construct five all-zero matrices of the same size as the existence probability maps, and add two-dimensional Gaussian functions at the key point coordinates of each matrix respectively to construct the ground truth;
[0046] Train the depth network parameters by reducing the loss between the existence probability map and the ground truth; where the expectation μ of the two-dimensional Gaussian function is 0, and the variance σ can be dynamically adjusted according to the size of the tracking target.
[0047] S103: Calculate the difference between the existence probability map and the ground truth using a binary classification loss function; train the deep network by means of Stochastic Gradient Descent (SGD), etc.; the binary classification loss function can be FocalLoss.
[0048] S104: Based on the trained network model, input the template frame and the tracking frame, output five key point existence probability maps, select the position with the maximum value in the existence probability map as the position of the key point, and output the coordinates of the five key points respectively, denoted as P i (x i ,y i ), i ∈ (A~E).
[0049] S2: Based on the coordinates of each key point, perform pairwise combinations to generate non-overlapping target boxes, and screen the generated target boxes;
[0050] In this embodiment, after generating the target boxes, it further includes calculating the pairwise intersection-over-union matrix of the target boxes, and obtaining the sum of the pairwise intersection-over-union ratios of each target box with the remaining target boxes for target box screening.
[0051] The specific process is as follows:
[0052] S201: Based on the output coordinates of the five key points P i (x i ,y i ), i ∈ (A~E), perform pairwise combinations on them to generate six non-overlapping target boxes B m (x m ,y m ,w m ,h m ), m ∈ {AD, BC, AE, BE, CE, DE}, where (x m ,y m ) represents the center of the bounding box, and (w m ,h m ) represents the width and height of the bounding box;
[0053] In this embodiment, the pairwise combinations are specifically: upper left point + lower right point, lower right point + upper left point, upper left point + center point, lower left point + center point, upper right point + center point, lower right point + center point.
[0054] S202: Calculate the pairwise intersection-over-union matrix of the six possible target boxes B m , sum the matrix horizontally to obtain the sum of the pairwise intersection-over-union ratios of each target box with the remaining target boxes for target box screening;
[0055] Although the six non-overlapping target bounding boxes contain rich information, they inevitably contain noise. Therefore, it is necessary to first screen the six possible target bounding boxes. The intersection over union (IoU) of two bounding boxes can reflect the similarity between them. The larger the IoU, the higher the similarity. The six possible target bounding boxes themselves contain a lot of redundant information. Ideally, the IoU of any two bounding boxes is 1. In this embodiment, the method of calculating the IoU matrix is used to calculate the pairwise IoU matrix of the six target bounding boxes, and the sum of the matrix horizontally is obtained to get the sum of the pairwise IoUs of each target bounding box and the other target bounding boxes. Among them, a smaller sum of the target bounding box IoUs indicates that the box is more different from the other boxes and has a greater possibility of containing noise, and it needs to be removed.
[0056] S203: Set the screening value l, select the target bounding boxes with the top l in terms of the sum of IoUs, and vote on the selected target bounding boxes;
[0057] Through experiments and visual analysis, it is found that the effect of retaining half of the bounding boxes is better. Therefore, the screening value l is set to half of the number of target bounding boxes obtained in step S201. Specifically, in this embodiment, the screening value l is 3, and the target bounding boxes with the top three in terms of the sum of IoUs are selected for the final vote.
[0058] S3: Calculate the final target bounding box according to the screened target bounding boxes through voting.
[0059] In this embodiment, the target bounding box voting is divided into two parts: calculating the voting weight of the target bounding box and target bounding box voting;
[0060] The specific process is as follows:
[0061] S301: Calculate the voting weight of each participating target bounding box. During the voting process, each target bounding box should have its own voting weight, and this weight should be related to the attributes of the target bounding box itself;
[0062] In this embodiment, starting from the key points that make up the target bounding box, the quality of the key point is defined as the value corresponding to the point in the probability map of existence. Then the score of each bounding box is the sum of the scores of the two points that make up the box. The quality of the normalized target bounding box is used as the voting weight of the corresponding target bounding box, denoted as V j , j ∈ (1, 3).
[0063] S302: Weight the target bounding boxes based on the target bounding box voting weight to obtain the final target bounding box, specifically as follows:
[0064]
[0065] Among them, (x j , y j ) represents the center of the bounding box V j , and (w j , hj ) represents the width and height of the bounding box V j .
[0066] Embodiment 2
[0067] Based on the above Embodiment 1, Embodiment 2 of the present invention discloses an object tracking system based on key point depth existence probability map voting, including a key point generation module, an object box generation module, and an object box voting module, which respectively store the algorithms of the corresponding steps in the above Embodiment 1.
[0068] The key point generation module obtains the existence probability map of key points based on the template frame and the tracking frame, and obtains the coordinates of five key points according to the existence probability map of key points;
[0069] Specifically, the key point generation module inputs the template frame and the tracking frame into a deep learning network model with five key point existence probability maps output by the head, obtains the existence probability map of key points, and constructs the ground truth by adding a two-dimensional Gaussian function at the key point coordinates of a all-zero matrix with the same size as the existence probability map. The parameters of the deep network model are trained by reducing the loss between the existence probability map and the ground truth. During tracking, the position of the maximum value of each existence probability map is selected as the position of the key point, and the coordinates of five key points are output;
[0070] The deep learning network model is based on a siamese network structure or a Transformer deep network structure, and the network head is operated by convolution stacking, upsampling, etc. to adjust the feature size to form the network model architecture;
[0071] The expectation μ of the two-dimensional Gaussian function is 0, and the variance σ can be dynamically adjusted according to the size of the tracking target;
[0072] The object box generation module is used to combine the coordinates of five key points in pairs to generate six non-overlapping object boxes, calculate the pairwise intersection-over-union matrix of the six non-overlapping object boxes, sum the matrix horizontally to obtain the sum of the pairwise intersection-over-union of each object box with the other object boxes, and perform screening based on the sum of the intersection-over-union;
[0073] The object box voting module is used to calculate the voting weight of the object boxes participating in the voting, and weight the object boxes based on the object box voting weight to obtain the final object bounding box.
[0074] The present invention is not limited to the foregoing specific embodiments. The present invention extends to any new feature or any new combination disclosed in this specification, as well as any new combination of the steps of any new method or process disclosed.
Claims
1. A target tracking method based on voting of the depth existence probability map of key points, characterized in that Including: S1: Obtain a template frame and a tracking frame, input them into a depth network model, generate five key-point feature maps, and output key-point coordinates. The specific process is as follows: S101: Construct a depth network model based on a Siamese or Transformer depth network structure. Through convolution stacking and upsampling operations on the head of the network model, adjust the feature size so that the head outputs five key-point existence probability maps. The five key-point existence probability maps respectively represent the existence probabilities of the upper-left point, lower-right point, lower-left point, upper-right point, and center point of the target. S102: Construct five all-zero matrices of the same size as the existence probability maps, and add two-dimensional Gaussian functions at the key-point coordinates of each matrix respectively to construct the ground truth. S103: Use a binary classification loss function to calculate the difference between the existence probability maps and the ground truth, and train the network model that outputs five key-point existence probability maps. S104: Based on the trained network model, input the template frame and the tracking frame, output five key-point existence probability maps, select the position with the maximum value in the existence probability maps as the position of the key point, and output five key-point coordinates respectively. S2: Based on each key-point coordinate, perform pairwise combinations to generate non-overlapping target boxes, calculate the pairwise intersection-over-union matrix of the target boxes, obtain the sum of the pairwise intersection-over-unions of each target box and the remaining target boxes, and screen the generated target boxes. The specific process is as follows: S201: Based on the five key point coordinates P of the output i (x i , y i ), i ∈ (A~E), generate six non - overlapping target bounding boxes B m (x m , y m , w m , h m ), m ∈ {AD, BC, AE, BE, CE, DE}, where (x m , y m ) represents the center of the bounding box, and (w m , h m ) represents the width and height of the bounding box; S202: Calculate the pairwise intersection-over-union matrix of six possible target bounding boxes B m and sum the matrix horizontally to obtain the sum of the pairwise intersection-over-union of each target bounding box with the rest of the target bounding boxes for target bounding box screening; S203: Set a screening value l, select the top l target boxes with the sum of intersection-over-unions, and vote on the selected target boxes. The screening value l is half of the number of target boxes obtained in step S201. S3: Perform voting calculations based on the screened target boxes to obtain the final target bounding box.
2. The target tracking method according to claim 1, characterized in that For step S3, the specific process is as follows: S301: Calculate the voting weight of each target box participating in the voting. The voting weight is related to the attributes of the target box itself. S302: Weight the target boxes based on the voting weights of the target boxes to obtain the final target bounding box.
3. The target tracking method according to claim 2, wherein For step S301, specifically: Define the quality of a key point as the value corresponding to the point in the existence probability map. Then the score of each box is the sum of the scores of the two points constituting the box, and use the quality of the normalized target box as the voting weight corresponding to the target box.
4. A target tracking system based on voting of key point depth existence probability maps, characterized in that, Including a key-point generation module, a target-box generation module, and a target-box voting module; The key-point generation module obtains the key-point existence probability maps based on the template frame and the tracking frame, and obtains five key-point coordinates according to the key-point existence probability maps. The specific process is as follows: S101: Construct a depth network model based on a Siamese or Transformer depth network structure. Through convolution stacking and upsampling operations on the head of the network model, adjust the feature size so that the head outputs five key-point existence probability maps. The five key-point existence probability maps respectively represent the existence probabilities of the upper-left point, lower-right point, lower-left point, upper-right point, and center point of the target. S102: Construct five all-zero matrices of the same size as the existence probability maps, and add two-dimensional Gaussian functions at the key-point coordinates of each matrix respectively to construct the ground truth. S103: Calculate the difference between the existence probability map and the ground truth using a binary classification loss function, and train a network model that outputs five key-point existence probability maps. S104: Based on the trained network model, input the template frame and the tracking frame, output five key-point existence probability maps, select the position with the maximum value in the existence probability map as the position of the key point, and output the coordinates of the five key points respectively. The target box generation module is used to combine the coordinates of the five key points in pairs to generate six non-overlapping target boxes, calculate the pairwise intersection-over-union matrix of the six non-overlapping target boxes, sum the matrix horizontally to obtain the sum of the pairwise intersection-over-union ratios of each target box with the other target boxes, and perform screening based on the sum of the intersection-over-union ratios. The specific process is as follows: S201: Based on the five key point coordinates P of the output i (x i ,y i ), i ∈ (A~E), generate six non-overlapping target bounding boxes B m (x m ,y m ,w m ,h m ), m ∈ {AD, BC, AE, BE, CE, DE}, where (x m ,y m ) represents the center of the bounding box, and (w m ,h m ) represents the width and height of the bounding box; S202: Calculate the pairwise intersection-over-union matrix of six possible target bounding boxes B m and sum it horizontally to obtain the sum of the pairwise intersection-over-union of each target bounding box with the rest of the target bounding boxes for target bounding box screening; S203: Set a screening value l, select the top l target boxes with the sum of the intersection-over-union ratios, and vote on the selected target boxes; the screening value l is half of the number of target boxes obtained in step S201. S3: Calculate the vote based on the screened target boxes to obtain the final target bounding box. The target box voting module is used to calculate the voting weights of the target boxes participating in the vote, and weight the target boxes based on the target box voting weights to obtain the final target bounding box.
Citation Information
Patent Citations
Attitude tracking method based on key point screening
CN113850221A
Target Detection Method, Training Method, Electronic Device, and Computer-Readable Medium
US20210256680A1