A planar object tracking method based on center point detection and graph matching

By adopting the method of center point detection and graph matching in plane target tracking, the problem of poor tracking effect of motion blur and unconstrained scenes in the prior art is solved, and higher tracking accuracy and robustness are achieved.

CN115239763BActive Publication Date: 2025-06-27BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210853244.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-20
Publication Date
2025-06-27
Estimated Expiration
2042-07-20

AI Technical Summary

Technical Problem

There is a lack of a key point-based depth plane target tracking method that is particularly effective for motion fuzzy and unconstrained scenarios in the prior art.

Method used

The method based on center point detection and graph matching is adopted to predict the center point of the target through the central positioning network, and a matching matrix of the template image and the target area is established using the depth map matching network, and the geometric transformation of the target is estimated in combination with the RANSAC algorithm.

Benefits of technology

The accuracy and robustness of tracking are significantly improved in scaling, rotation, perspective transformation, motion blur, partial occlusion and unconstrained scenarios, especially in partial occlusion, motion blur and unconstrained scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115239763B_ABST
    Figure CN115239763B_ABST
Patent Text Reader

Abstract

The present invention provides a planar object tracking method based on center point detection and graph matching. The method includes: using a center localization network to predict the center point of the tracking object in the current frame, and determining an initial target region according to the predicted center point; modeling the template image and the target region as a complete graph composed of two sub-graphs, the two sub-graphs corresponding to two regions respectively, and using a deep graph matching network to predict the matching matrix between the template image and the target region; using the RANSAC algorithm to estimate the geometric transformation of the target from the template image to the current image from the matching pairs identified by the matching matrix, so as to obtain the predicted position of the tracking object. The method of the present invention generally performs better than previous methods in scenarios of scaling, rotation, perspective transformation, motion blur, partial occlusion and unconstrained scenarios, and particularly achieves greater gains in scenarios of partial occlusion, motion blur and unconstrained scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and in particular, to a planar object tracking method based on center point detection and graph matching. Background Art

[0002] With the continuous development of Internet technology and computer vision, the perception and analysis of image information have provided great convenience for people's lives. Nowadays, the channels for obtaining image information are diverse and extensive, such as mobile phones, surveillance cameras, cameras, and so on. The influx of a large amount of visual information also urgently requires efficient methods or technologies to process it. Among these technologies, planar object tracking plays an important role in tracking specific objects. Planar object tracking is widely used in many vision-based robot applications and related fields, such as visual SLAM (Simultaneous Localization and Mapping) and augmented reality. It provides strong support for the tracking or modeling of three-dimensional objects through geometric transformations obtained from the analysis of two-dimensional images.

[0003] For a video sequence, the purpose of planar object tracking is to estimate the position where the target appears in subsequent frames given the planar object to be tracked in the initial frame. This problem is usually transformed into estimating a 2D geometric transformation, such as an affine transformation or a perspective transformation (also known as a homography). Although there have been some excellent works proposed in this field, they still perform poorly when the target is occluded and moving rapidly.

[0004] Currently, the planar object tracking methods in the prior art can be divided into two categories: template-based methods and key-point-based methods. The essence of template-based methods is to solve an optimization problem. It minimizes the difference between the template and the search area to track the object, such as the ESM (Efficient Second-Order Minimization) algorithm. Key-point-based methods usually model the template and the search area as two sets of key points, then establish the correspondence between them, and finally use geometric verification methods to estimate the 2D transformation, such as the Gracker algorithm. Compared with template-based methods, key-point-based methods have a natural advantage in dealing with partial occlusion.

[0005] In recent years, the development of deep learning has injected new vitality into this field and gradually become a research hotspot in the direction of planar object tracking. These deep methods are usually designed based on template methods. They regress and predict the position or homography of the target by fusing the deep global features of the template and the search image, such as the HomographyNet algorithm and the HDN algorithm. In addition, some methods are dedicated to constructing descriptors that are robust to geometric transformations to improve the accuracy of tracking, such as the GIFT operator and the LISRD operator. However, few deep learning methods unify the acquisition and matching of key points into a complete framework.

[0006] Currently, there is no key-point-based deep planar object tracking method in the prior art that is particularly effective for motion blur and unconstrained scenarios. Summary of the Invention

[0007] An embodiment of the present invention provides a planar object tracking method based on center point detection and graph matching to effectively track planar objects in images.

[0008] To achieve the above object, the present invention adopts the following technical solutions.

[0009] A planar object tracking method based on center point detection and graph matching includes:

[0010] Using a center localization network to predict the center point of the tracking target in the current frame, and determining the initial target region according to the predicted center point;

[0011] Modeling the template image and the target region as a complete graph composed of two subgraphs, the two subgraphs corresponding to the two regions respectively, and using a deep graph matching network to predict the matching matrix between the template image and the target region;

[0012] Using the RANSAC algorithm to estimate the geometric transformation of the target from the template image to the current image from the matching pairs identified by the matching matrix, and obtaining the predicted position of the tracking target in the current frame.

[0013] Preferably, before using the center localization network to predict the center point of the tracking target in the current frame, training the center localization network is also included, and the training process includes:

[0014] Step 1: Obtain the training part of the publicly available tracking picture dataset. During data preprocessing, a region in the template image that is 2 2 times larger than the target is used as the template region, and a region in the search image that is 5 2An area that is several times larger is used as the search area, and the sizes are scaled to 128×128 and 320×320 respectively. That is, the formats of the input images are [C1, H1, W1] = [3, 128, 128] and [C2, H2, W2] = [3, 320, 320], where C represents the number of channels, H represents the height of the image, and W represents the width of the image;

[0015] Step 2: Use ResNet50 with the last layer4 layer, pooling layer, and FC layer removed as the backbone network to extract features from the images. The dimension of the extracted features is 1024. Use a 1×1 convolutional kernel to reduce the dimension of the extracted features. The dimension of the features after dimension reduction is 256. Use non-learnable sine-cosine encoding to perform position encoding on the elements at each position in the feature map;

[0016] Step 3: Flatten the two parts of the feature vectors and concatenate them along the spatial dimension to obtain the feature vector and send it into the Encoder module, where d = 256. The Encoder encoder enhances the original features through self-attention and cross-attention and captures the corresponding relationships between them, obtaining the ability to distinguish the spatial position where the target is located. The feature vector of the search area after Encoder encoding and the query vector q ∈ R 1×d Do attention again to decode the position information of the target in the search area, as shown in formula (1) specifically:

[0017]

[0018] The decoded information is The feature vector f′ x Is transformed through dimension transformation into And is sent to a stacked fully convolutional network. After passing through the fully convolutional network, the channel dimension of f is reduced to 1, obtaining a probability map for predicting the center point position Calculate the expected value of the probability map distribution in the grid coordinate space to obtain the predicted target center point, as shown in formula (2) specifically:

[0019]

[0020] Step 4: Use l1 loss as the loss function for training. The specific formula is as shown in (3), where and c i =(c x , c y ) represent the predicted target center point and the true target center point label respectively. Use AdamW as the optimizer to optimize the parameters of the network model according to the loss value:

[0021]

[0022] Obtain a trained center localization network.

[0023] Preferably, using the center localization network to predict the center point of the tracking target in the current frame, and determining the initial target area according to the predicted center point, includes:

[0024] Input consecutive video frames into the center localization network, where the first frame is the template frame, and the area where the target is located is called the template. Store the features extracted by ResNet50 in the center localization network for the template area to avoid repeated calculations. The template area refers to the area with the center of the template as the center point, and the width and height are 2 times the width and height of the template respectively. Use the position offset of the template in the first frame as the initial motion parameter;

[0025] During tracking, first perform an inverse transformation on the currently read image using the motion parameters tracked in the previous frame to obtain a resampled image. At the same time, the position tracked at the previous moment will correspond to a quadrilateral area in the resampled image. Use the center of this quadrilateral as the center point, and 5 times the width and height of the template as the size to crop, fill, and scale the resampled image to obtain a search area. Send the template area and the search area into the center localization network to obtain the predicted position (c x , c y ) of the target center point. Crop an area with the same size as the template centered at (c x , c y ) as the located initial target area.

[0026] Preferably, before using the depth map matching network to predict the matching matrix between the template image and the target area, it also includes training the depth map matching network. The training process includes:

[0027] Step 1: Obtain a public graph matching dataset. The dataset includes a template image (P), a search image (Q), key points in the template image and their descriptors key points in the search image and their descriptors and the correspondence relationship (M) between the key points in the template image and the key points in the search image. During data preprocessing, uniformly adjust the size of the images to 256×256, that is, the size of the input image is [C, H, W] = [3, 256, 256], where C represents the number of channels, H represents the height of the image, and W represents the width of the image;

[0028] Step 2: Model the template image P and the search image Q as graphs according to the Delaunay triangulation algorithm, which are respectively represented as and where Represents a vertex, Represents an edge, Represents the characteristics of a vertex, ε represents the characteristics of an edge. According to the feature similarity between two point sets ( and ), cross edges are constructed to connect the two subgraphs into a complete graph Specifically, for any In , select the top-k points that are most similar in appearance to point v to build edges, where the appearance similarity is defined as:

[0029] Step 3: The graph first aggregates along all edges and updates the node information, as shown in formulas (4) and (5):

[0030]

[0031]

[0032] Among them, N(v) represents the neighbors of node v, M V Represents the node information aggregation function, and respectively represent the information of node w and edge v→w at the t-th transmission, Represents the neighbor information of node v at the (t + 1) th transmission, U V Represents the node update function, which aggregates the information of the edges connected to each node to obtain the neighbor information, and then fuses and updates the neighbor information and the original node information as the new state of the node;

[0033] After updating the node state, the graph will perform edge state update, which is also divided into two steps: aggregation and update, as shown in formulas (6) and (7) specifically:

[0034]

[0035]

[0036] Among them, respectively represent the feature vectors of the source node and the destination node of edge v→w at the t-th transmission, M E Represents the edge information transmission function, Represents the state of edge v→w at the t-th transmission, Represents the neighbor information of this edge at the (t + 1) th transmission, U EIt represents an edge update function that aggregates the information of the source node and the destination node of each edge to obtain neighbor information, and then fuses and updates the neighbor information with the original edge information as the new state of the edge;

[0037] Step 4: Use the weighted l2 loss as the loss function. The specific formula is shown in (9), where λ is a hyperparameter for balancing positive and negative samples, S and M respectively represent the predicted score matrix and the true matching matrix label, and Adam is used as the optimizer. The parameters of the network model are optimized by backpropagation according to the loss value:

[0038]

[0039]

[0040] Preferably, modeling the template image and the target region as a complete graph composed of two subgraphs, the two subgraphs corresponding to the two regions respectively, and using a deep graph matching network to predict the matching matrix of the template image and the target region includes:

[0041] Use the SuperPoint network to extract feature points from the target region, model the extracted feature points as a graph using the Delaunay triangulation algorithm, merge this graph with the subgraph corresponding to the template image to model a complete graph, send the complete graph into the graph matching network for information transmission, aggregation and update. After the state of the graph is updated, use a linear layer to estimate the matching confidence of the corresponding nodes from the features on the cross edges to obtain a score matrix, and use the greedy algorithm to perform 0 / 1 processing on the score matrix to obtain the matching matrix of the template image and the target region, and the value of this matching matrix is 0 or 1.

[0042] Preferably, using the RANSAC algorithm to estimate the geometric transformation of the target from the template image to the current image from the matching pairs identified by the matching matrix to obtain the predicted position of the tracked target in the current frame includes:

[0043] Obtain the matching pairs of feature points in the template image and the target region according to the matching matrix, use the RANSAC algorithm to filter out the outliers in the matching pairs, estimate the transformation matrix using the remaining feature point matching pairs, and perform a geometric transformation from the template image to the current image on the initial position of the template using this transformation matrix to obtain the predicted position of the tracked target in the current frame;

[0044] If the number of elements with a confidence level higher than 0.9 in the score matrix is less than 4, it is considered that the target is lost, and the relocalization mechanism is started. The search area is directly determined on the current frame image according to the position tracked in the previous frame, and the search area is input into the center localization network and the subsequent steps are executed.

[0045] As can be seen from the technical solutions provided by the embodiments of the present invention described above, the performance of the method of the present invention has been improved in scaling, rotation, perspective transformation, motion blur, partial occlusion, and unconstrained scenarios, and particularly significant benefits have been obtained in partial occlusion, motion blur, and unconstrained scenarios.

[0046] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0048] Figure 1 FIG. is a schematic diagram of the implementation principle of a planar target tracking method based on center point detection and graph matching provided by an embodiment of the present invention.

[0049] Figure 2 FIG. is a schematic structural diagram of a center positioning network provided by an embodiment of the present invention.

[0050] Figure 3 FIG. is a schematic structural diagram of a depth graph matching network provided by an embodiment of the present invention.

[0051] Figure 4 FIG. is a training flowchart of a center positioning network provided by an embodiment of the present invention.

[0052] Figure 5 FIG. is a training flowchart of a graph matching network provided by an embodiment of the present invention. Figure 6 FIG. is a processing flowchart of a planar target tracking method based on center point detection and graph matching provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.

[0054] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any and all combinations of any one of the one or more associated listed items.

[0055] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.

[0056] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments as examples in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.

[0057] The embodiments of the present invention propose a planar target tracking method that is more robust to different motion states of the target. The performance of this method in fast motion and motion blur scenarios has been greatly improved. This method combines deep learning with a key-point-based tracking method, and proposes a planar target tracking method based on center point detection and graph matching, unifying the acquisition and matching of key points into a complete framework.

[0058] The implementation principle of the planar target tracking method based on center point detection and graph matching in the embodiments of the present invention is as Figure 1 shown. This method first predicts the center point of the tracking target in the current frame through a center localization network, and determines the initial target region according to the predicted center point. Intuitively, this gives a reliable guess of the initial position of the tracking target and creates good initial conditions for the matching stage. Then, the template image and the target region are modeled as a complete graph composed of two subgraphs, and the two subgraphs correspond to the two regions respectively. After that, the correspondence between the two point sets is established through a deep graph matching network. Finally, the homography is calculated from the matched key point pairs through the RANSAC (Random Sample Consensus) algorithm.

[0059] The process of training a center positioning network provided by an embodiment of the present invention is as follows Figure 2 shown, including the following processing procedures:

[0060] Step 1: Obtain the training parts of the publicly available tracking image datasets MSCOCO2017, GOT-10K, and LaSOT. During data preprocessing, the area in the template image that is 2 2 times larger than the target is used as the template area, and the area in the search image that is 5 2 times larger than the target is used as the search area, and they are respectively scaled to 128×128 and 320×320, that is, the formats of the input images are [C1, H1, W1] = [3, 128, 128], [C2, H2, W2] = [3, 320, 320], where C represents the channel, H represents the height of the image, and W represents the width of the image. Some data augmentation means are used for the search area, including horizontal flipping, brightness perturbation, and central perturbation.

[0061] Step 2: Use ResNet50 with the last layer4 layer, pooling layer, and FC layer removed as the backbone to extract features from the image, and the extracted feature dimension is 1024. Then, a 1×1 convolutional kernel is used to reduce the dimension of the extracted features, and the dimension of the reduced features is 256. After that, non-learnable sine-cosine encoding is used to perform position encoding on the elements at each position in the feature map.

[0062] Step 3: Flatten the two parts of the feature vectors and splice them together along the spatial dimension to obtain the feature vector and send it into the Encoder module, where d = 256. The Encoder encoder will enhance the original features through self-attention and cross-attention and capture the corresponding relationships between them, so as to obtain the ability to distinguish the spatial position where the target is located. The feature vector of the search area after being encoded by the Encoder will do another attention with the query vector q∈R 1×d to decode the position information of the target in the search area, as specifically shown in formula (1).

[0063]

[0064] Here, the decoded information Then, the feature vector f′ x will be transformed through dimensional transformation into and sent to a stacked fully convolutional network. After passing through the fully convolutional network, the channel dimension of f is reduced to 1, and a probability map for predicting the center point position is obtained We calculate the expected value of the probability map distribution in the grid coordinate space to obtain the predicted target center point, as specifically shown in formula (2).

[0065]

[0066] Step 4: Use l1 loss as the loss function for training. The specific formula is as shown in (3), where and represent the predicted target center point and the true target center point label respectively. Use AdamW as the optimizer to optimize the parameters of the network model according to the loss value.

[0067]

[0068] The process of training a depth map matching network provided by an embodiment of the present invention is as Figure 3 shown, including the following processing procedures:

[0069] Step 1: Obtain a public graph matching dataset. The dataset should include a template image (P), a search image (Q), the key points in the template image and their descriptors the key points in the search image

[0070] and their descriptors and The corresponding relationship (M) between the key points in the template image and the key points in the search image. During data preprocessing, uniformly adjust the size of the image to 256×256, that is, the size of the input image is [C, H, W] = [3, 256, 256], where C represents the number of channels, H represents the height of the image, and W represents the width of the image.

[0070] Step 2: Model the template image P and the search image Q as graphs according to the Delaunay triangulation algorithm, which are respectively represented as and where represents vertices, represents edges, represents the features of vertices, and ε represents the features of edges. Then, construct cross edges according to the feature similarity between the two point sets ( and ) to connect the two subgraphs into a complete graph Specifically, for any we select the top-k points in that are most similar in appearance to the point to build edges. The appearance similarity is defined as:

[0071] Step 3: The graph first aggregates along all edges and updates the node information, as shown in formulas (4) and (5).

[0072]

[0073]

[0074] Among them, N(v) represents the neighbors of node v, and Mv represents the node information aggregation function. and respectively represent the information of node w and edge v→w at the t-th transmission. represents the neighbor information of node v at the (t + 1) th transmission, and U V represents the node update function. Briefly speaking, it is to aggregate the information of the edges connected to each node to obtain the neighbor information, and then fuse and update the neighbor information and the original node information as the new state of the node.

[0075] After updating the node state, the graph will update the edge state, which is also divided into two steps: aggregation and update, as specifically shown in formulas (6) and (7).

[0076]

[0077]

[0078] Among them, are respectively the feature vectors of the source node and the destination node of edge v→w at the t-th transmission, and M E represents the edge information transmission function. represents the state of edge v→w at the t-th transmission. represents the neighbor information of this edge at the (t + 1) th transmission, and U E represents the edge update function. Briefly speaking, it is to aggregate the information of the source node and the destination node of each edge to obtain the neighbor information, and then fuse and update the neighbor information and the original edge information as the new state of the edge. The aggregation functions (MV, ME) and update functions (UV, UE) in formulas (4)-(7) are both implemented by MLP, and this MLP includes a linear layer, a ReLU layer, and a LayerNorm layer. After the state of the graph is updated, we use a linear layer to estimate the matching confidence of the corresponding nodes from the features on the cross edges to obtain the score matrix S.

[0079] Step 4: Use the weighted l2 loss as the loss function, and the specific formula is as shown in (9), where λ is a hyperparameter for balancing positive and negative samples and is set to 50 in the experiment, and S and M respectively represent the predicted score matrix and the true matching matrix label. Adam is used as the optimizer, and the parameters of the network model are optimized by backpropagation according to the loss value.

[0080]

[0081]

[0082] The training flow chart of a center localization network provided by an embodiment of the present invention is as Figure 4 shown. The training flow chart of a graph matching network provided by an embodiment of the present invention is as Figure 5 shown. The processing flow chart of a planar target tracking method based on center point detection and graph matching provided by an embodiment of the present invention is as Figure 6 shown, including the following processing procedures:

[0083] Step 1: Data input and tracker initialization. The input data of the model is continuous video frames, where the first frame is the template frame, and the area where the target is located is called the template. To avoid repeatedly calculating the template during forward propagation, we store the features extracted by ResNet50 in the center localization network for the template area and the sub-graph corresponding to the template in the graph matching network. The template area refers to the area with the center of the template as the center point, and the width and height are 2 times the width and height of the template respectively. At the same time, we use the position offset of the template in the first frame as the initial motion parameter.

[0084] Step 2: Center localization network processing. In this stage, we first perform an inverse transformation on the read image using the motion parameters tracked in the previous frame to obtain a resampled image. Then, the position tracked at the previous moment also corresponds to a quadrilateral area in the resampled image. We use the center of this quadrilateral as the center point, and 5 times the width and height of the template as the size to crop, pad, and scale the resampled image to obtain a search area. Then, the template area and the search area are sent into the center localization network to obtain the predicted target center point position (c x , c y ). Finally, a region with the same size as the template is cropped with (c x , c y ) as the center as the located target area.

[0085] Step 3: Graph matching network processing. In Step 2, we obtained an initial rectangular target area. We use SuperPoint to extract feature points from it and use the Delaunay triangulation algorithm to model the extracted feature points as a graph. Then, it is merged with the sub-graph corresponding to the template image in Step 1 to model a complete graph. The constructed graph is sent into the graph matching network for information transfer, aggregation, and update. After the state of the graph is updated, we use a linear layer to estimate the matching confidence of the corresponding nodes from the features on the cross edges to obtain a score matrix. Finally, we use the greedy algorithm to process the score matrix to obtain a matching matrix with values of 0 or 1.

[0086] Step 4: From the matching matrix in Step 3, the corresponding relationship between the feature points in the template image and the target region can be obtained. We use the RANSAC algorithm to filter out the outliers that bring large differences, and then use the remaining feature point matches to estimate a transformation matrix. Use this transformation matrix to perform a geometric transformation on the initial position of the template to obtain the predicted position of the target in the current frame.

[0087] Step 5: Loss handling. In target tracking, target loss is a common situation. To improve the tracking accuracy, we added a loss detection and repositioning mechanism to the tracking framework. If the number of elements in the score matrix with a confidence level higher than 0.9 is less than 4, we consider that target loss has occurred because at least 4 pairs of matching feature points are required to calculate the perspective transformation. When target loss occurs, we will activate the repositioning mechanism. When target loss occurs, the motion parameters of the previous frame are generally not reliable. So instead of performing an inverse transformation on the current frame image using the motion parameters of the previous frame, we directly determine the search region on the current frame image based on the position tracked in the previous frame and input it into the center positioning network to perform subsequent steps.

[0088] Experiments and comparative analyses were carried out on the proposed planar target tracking method based on center point detection and graph matching and other advanced algorithms on the public dataset POT-210, which proved the effectiveness of the method proposed in the present invention. The results show that the method proposed in the present invention has better robustness to different motion states of the target, and has a leading advantage in dealing with partial occlusion, motion blur and unconstrained scenarios, achieving more accurate tracking.

[0089] Table 1. Comparison with other methods on the POT-210 dataset

[0090]

[0091] In summary, the embodiment of the present invention proposes a planar target tracking method that is more robust to the motion state of the target, and the performance of the proposed method in fast motion and motion blur scenarios has been greatly improved. Experimental data show that the tracking method proposed in the present invention has improved performance in scaling, rotation, perspective transformation, motion blur, partial occlusion and unconstrained scenarios, and has obtained greater benefits especially in partial occlusion, motion blur and unconstrained scenarios.

[0092] The present invention decomposes the planar target tracking task into two steps, that is, first predicting the coarse-grained initial target region of the tracking object, and then using the graph matching network to perform "refinement" to obtain the accurate position of the target, improving the tracking accuracy.

[0093] The central positioning network proposed by the present invention can first locate the initial position of the target when there is a large position offset of the tracked target. This can effectively reduce the search space of the model with little increase in computational complexity, reduce the occurrence of target loss, and is more robust than directly estimating the final position of the target. The graph matching network proposed by the present invention models the problem representation as a graph. Since the graph structure maintains a certain structural invariance in consecutive frames, the tracking stability is improved. Generally speaking, the two-stage tracking strategy proposed by the present invention achieves a more robust tracking effect for different motion states of the target through pre-positioning and graph matching techniques, and especially obtains better performance in large-scale motion and unconstrained scenarios. Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0094] From the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0095] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0096] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A planar target tracking method based on center point detection and graph matching, characterized in that Including: Predict the center point of the tracking target in the current frame using a center localization network, and determine the initial target region according to the predicted center point; Model the template image and the target region as a complete graph composed of two subgraphs, where the two subgraphs correspond to two regions respectively, and use a depth graph matching network to predict the matching matrix between the template image and the target region; Use the RANSAC algorithm to estimate the geometric transformation of the target from the template image to the current image from the matching pairs identified by the matching matrix, and obtain the predicted position of the tracking target in the current frame; Before using the depth graph matching network to predict the matching matrix between the template image and the target region, it also includes training the depth graph matching network. The training process includes: Step 1: Obtain a publicly available image matching dataset, which includes a template image (P), a search image (Q), key points in the template image and their descriptors (v P ), key points in the search image and their descriptors (v Q ), and the correspondence (M) between the key points in the template image and the key points in the search image. During data preprocessing, the size of the images is uniformly adjusted to 256×256, that is, the size of the input image is [C, H, W] = [3, 256, 256], where C represents the number of channels, H represents the height of the image, and W represents the width of the image; Step 2: Model the template image P and the search image Q as graphs according to the Delaunay triangulation algorithm, denoted as and where represents vertices, represents edges, v represents the features of vertices, ε represents the features of edges. Cross edges are constructed according to the feature similarity between two point sets ( and ) to connect the two subgraphs into a complete graph Specifically, for any in , select the top-k points that are most similar in appearance to the point v to build edges, where the appearance similarity is defined as: Step three: The graph first aggregates along all edges and updates the node information, as shown in formulas (4) and (5): Among them, N(v) represents the neighbors of node v, and M V represents the node information aggregation function, and respectively represent the information of node w and edge v→w at the t-th transmission, represents the neighbor information of node v at the (t + 1) th transmission, and U V represents the node update function, which aggregates the information of the edges connected to each node to obtain the neighbor information, and then fuses the neighbor information with the original node information for update as the new state of the node; After updating the node state, the graph will update the edge state, which is also divided into two steps: aggregation and update, as shown in formulas (6) and (7) specifically: Among them, respectively represent the feature vectors of the source node and the destination node of the edge v→w at the t-th transmission, M E represents the edge information transmission function, represents the state of the edge v→w at the t-th transmission, represents the neighbor information of this edge at the (t + 1) th transmission, U E represents the edge update function, which aggregates the information of the source node and the destination node of each edge to obtain neighbor information, and then fuses and updates the neighbor information and the original edge information as the new state of the edge; Step four: Use the weighted l2 loss as the loss function, and the specific formula is as shown in (9), where λ is a hyperparameter for balancing positive and negative samples, S and M respectively represent the predicted score matrix and the true matching matrix label, use Adam as the optimizer, and backpropagate according to the loss value to optimize the parameters of the network model: The step of modeling the template image and the target region as a complete graph composed of two subgraphs, where the two subgraphs correspond to two regions respectively, and using a depth graph matching network to predict the matching matrix between the template image and the target region includes: Use the SuperPoint network to extract feature points from the target region, use the Delaunay triangulation algorithm to model the extracted feature points as a graph, merge this graph with the subgraph corresponding to the template image to form a complete graph, send the complete graph into the graph matching network for information transfer, aggregation and update. After the state of the graph is updated, use a linear layer to estimate the matching confidence of the corresponding nodes from the features on the cross edges to obtain a score matrix, and use the greedy algorithm to perform 0 / 1 processing on the score matrix to obtain the matching matrix between the template image and the target region, and the value of this matching matrix is 0 or 1.

2. The method according to claim 1, wherein Before using the center localization network to predict the center point of the tracking target in the current frame, it also includes training the center localization network. The training process includes: Step 1: Obtain the training part of the public tracking image dataset. During data preprocessing, the area in the template image that is 2 2 times larger than the target is used as the template area, and the area in the search image that is 5 2 times larger than the target is used as the search area, and the sizes are scaled to 128×128 and 320×320 respectively, that is, the formats of the input images are [C1, H1, W1] = [3, 128, 128], [C2, H2, W2] = [3, 320, 320], where C represents the number of channels, H represents the height of the image, and W represents the width of the image; Step two: Use ResNet50 with the last layer4 layer, pooling layer and FC layer removed as the backbone network to extract features from the image. The extracted feature dimension is 1024, use a 1×1 convolutional kernel to reduce the dimension of the extracted features, and the dimension of the reduced features is 256. Use non-learnable sine and cosine encoding to perform position encoding on the elements at each position in the feature map; Step 3: Flatten the two parts of the feature vectors and concatenate them along the spatial dimension to obtain the feature vector and send it into the Encoder module, where d = 256. The Encoder enhances the original features through self-attention and cross-attention and captures the corresponding relationships between them, obtaining the ability to discriminate the spatial position of the target. The feature vector of the search area after Encoder encoding and the query vector q ∈ R 1×d Perform attention again to decode the position information of the target in the search area, as shown in formula (1) specifically: The decoded information is Feature vector f′ x is transformed through dimensionality transformation into and is sent to a stacked fully convolutional network. After passing through the fully convolutional network, the channel dimension of f is reduced to 1, obtaining a probability map for predicting the center point position The expected value of the probability map distribution is calculated in the grid coordinate space to obtain the predicted target center point, as specifically shown in Equation (2): Step 4: Use l1 loss as the loss function for training. The specific formula is shown in (3), where and c i = (c x , c y ) represent the predicted target center point and the true target center point label respectively. Use AdamW as the optimizer to optimize the parameters of the network model according to the loss value: Obtain the trained center localization network.

3. The method according to claim 1, characterized in that, The step of using the center localization network to predict the center point of the tracking target in the current frame and determining the initial target region according to the predicted center point includes: Input consecutive video frames into the central localization network, where the first frame is the template frame, and the area where the target is located is called the template. Store the features extracted by ResNet50 in the central localization network for the template area to avoid repeated calculations. The template area refers to the area with the center of the template as the center point, and the width and height are 2 times the width and height of the template respectively. Use the position offset of the template in the first frame as the initial motion parameter; During tracking, first, an inverse transformation is performed on the currently read image using the motion parameters tracked in the previous frame to obtain a resampled image. At the same time, the position tracked at the previous moment corresponds to a quadrilateral region in the resampled image. Taking the center of this quadrilateral as the center point, and using 5 times the width and height of the template as the size, the resampled image is cropped, filled, and scaled to obtain a search region. The template region and the search region are sent to the center localization network to obtain the predicted target center point position (c x , c y ). Using (c x , c y ) as the center, a region with the same size as the template is cropped as the initially located target region.

4. The method according to claim 1, wherein The use of the RANSAC algorithm to estimate the geometric transformation of the target from the template image to the current image from the matching pairs identified by the matching matrix, and obtain the predicted position of the tracked target in the current frame, includes: Obtain the matching pairs of feature points in the template image and the target area according to the matching matrix, use the RANSAC algorithm to filter out the outliers in the matching pairs, estimate the transformation matrix using the remaining feature point matching pairs, and use this transformation matrix to perform a geometric transformation of the initial position of the template from the template image to the current image to obtain the predicted position of the tracked target in the current frame; If the number of elements in the score matrix with a confidence level higher than 0.9 is less than 4, it is considered that the target is lost. Start the relocalization mechanism, directly determine the search area on the current frame image according to the position tracked in the previous frame, input the search area into the central localization network and execute the subsequent steps.