A multi-target tracking method based on interactive spatio-temporal feature of graph neural network
Patent Information
- Application Number
- CN202311765334.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-19
AI Technical Summary
[0005]针对现有技术中的上述不足,本发明提供的一种基于交互时空特征的图神经网络的多目标跟踪方法解决了现有的基于孪生网络的跟踪方法进行多目标跟踪的准确性不高的问题
[0047](1)本发明提供了一种基于交互时空特征的图神经网络的多目标跟踪方法,将时空图注意力网和上下文特征网络加入网络架构中,可以将目标的时空外观建模和目标间的交互以及上下文引导下的自适应学习结合起来,实现鲁棒目标定位与跟踪,还根据不同目标之间的关联程度不同这一特性,利用注意力机制使得模型在处理图数据时关注重要的节点和边,有效的提高跟踪的准确性。
Smart Images

Figure CN117830352B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision, specifically relating to a multi-target tracking method based on graph neural networks with interactive spatiotemporal features. Background Technology
[0002] Multi-object tracking is a key task in computer vision, widely used in autonomous driving, intelligent surveillance, and behavior recognition. This task begins with a continuous sequence of video frames as input, and after processing the video frames, outputs the tracking results of the targets, including their bounding boxes and corresponding ID numbers. Traditional multi-object tracking methods, such as those based on probability density distributions and particle filtering, cannot handle or adapt to complex tracking variations; their robustness and accuracy have been surpassed by cutting-edge algorithms. With the development of deep learning, deep learning technology has also been successfully applied to the tracking field. In the context of big data, training network models using deep learning yields more expressive convolutional features, resulting in better tracking results.
[0003] In recent years, Siamese network-based tracking has attracted increasing attention in the tracking community. Siamese networks learn a similarity metric between a target object and candidate patches in the current search image within an end-to-end framework. Leveraging powerful deep networks and large-scale labeled video frames for offline training, Siamese network-based trackers achieve good performance and efficiency. This network transforms the tracking task into a template matching problem rather than a common binary classification problem. Siamese networks are a special structure within convolutional neural networks, consisting of two identical sub-networks. The network input consists of two images: a template image (usually the first frame of the sequence) and a search image (a subsequent frame). Each sub-network processes one image, extracting features through forward computation. These features are then passed through a similarity metric function to calculate a heatmap representing the similarity between each location in the search image and the template image. The entire tracking process does not require template updates, significantly improving the algorithm's speed. This is a key challenge that neural networks in deep learning have long struggled to achieve in target tracking. It wasn't until the application of Siamese networks to target tracking that good real-time performance was achieved while maintaining high tracking accuracy, marking a significant breakthrough in deep learning. Following this network, numerous target tracking algorithms based on it have emerged, achieving a balance between accuracy and speed, and securing a crucial position for deep learning applications in target tracking. Compared to other tracking methods, Siamese network-based tracking has a clear advantage when facing challenges such as real-time performance, small-range object movement, and motion blur. However, when the target undergoes significant deformation, a large difference can occur between the target candidate box and the target template, leading to tracking failure. Furthermore, if the target is occluded, moves rapidly, or has a similar appearance, the search image size may not be large enough to cover the target, resulting in an incorrect similarity metric. As errors accumulate during tracking, the tracking becomes irreversible. Therefore, the tracking performance of Siamese networks degrades in complex backgrounds.
[0004] The reason for the above problems is that most existing Siamese network methods do not fully utilize the spatiotemporal appearance modeling of targets in different contexts and the interactions between targets. Existing Siamese network-based tracking methods mostly ignore the contextual information of the search image to guide the adaptation of the target appearance model. Due to the lack of online adaptability, they struggle to capture changes in the target object, background, or context in the search image, potentially leading to tracking failures. Summary of the Invention
[0005] In view of the above-mentioned shortcomings in the prior art, the present invention provides a multi-target tracking method based on interactive spatiotemporal features of graph neural networks, which solves the problem of low accuracy in multi-target tracking of existing tracking methods based on Siamese networks.
[0006] To achieve the aforementioned objectives, the present invention employs the following technical solution: a multi-target tracking method based on a graph neural network with interactive spatiotemporal features, comprising the following steps:
[0007] S1. Obtain historical sample images and the current search image, and input the historical sample images and the current search image into the feature extraction network respectively to obtain SFP features and SFC features; wherein, the feature extraction network includes an interconnected shared convolutional network and a spatiotemporal graph attention network;
[0008] S2. Input the current search image into a shared convolutional network, a max pooling layer, and a deconvolutional layer in sequence to obtain contextual features;
[0009] S3. Adaptive features are generated based on SFP features, SFC features and context features through the context feature network to obtain the predicted position of each target in the current search image, thus completing multi-target tracking.
[0010] Further: In S1, the feature extraction network processes historical sample images and the current search image using the same method, wherein the specific method for processing historical sample images is as follows:
[0011] S11. Input the historical sample images into the shared convolutional network to obtain the feature set;
[0012] S12. Construct an undirected graph based on the feature set, input the undirected graph into a spatiotemporal graph attention network to obtain a refined feature set, and aggregate the refined feature set to obtain SFP features.
[0013] Furthermore: In S11, the historical sample images include several consecutive video frames;
[0014] The specific method for obtaining the feature set is as follows:
[0015] Based on several consecutive video frames and the target detection features in each frame, a feature set is obtained by embedding features of the input target object into different target regions in all video frames through a shared convolutional network.
[0016] Further: In S12, the undirected graph includes a node set and an edge set. The node set is composed of all target parts in the embedding sequence of the feature set. The edge set includes first and second type edges. The first type of edge is a spatial edge, which represents the intra-sample connection in each frame. The method to obtain the second type of edge is to connect the parts in consecutive frames that are at the same position as the temporal edge.
[0017] Furthermore: In S12, the refined feature set includes vertex features and neighbor features, and the method for obtaining SFP features is as follows:
[0018] S121. Calculate vertex features Characteristics of neighbors similarity coefficient e ij ;
[0019]
[0020] In the formula, W represents the shared parameters, || represents the feature concatenation operation, and a(·) represents the feature mapping operation. Let j be the set of node features adjacent to vertex i, where j is the index of the node feature;
[0021] S122. Normalize the attention coefficients based on the similarity coefficients to obtain the normalized attention coefficients α. ij ;
[0022]
[0023] In the formula, e ik Let be the attention coefficient, LeakerReLU(·) be the activation function, and k be the index of the neighboring node of vertex i.
[0024] S123. The node features to be aggregated are weighted and summed according to the normalized attention coefficients, and then enhanced through a multi-head attention mechanism to obtain enhanced new node features.
[0025]
[0026] In the formula, || represents the connection operation, K is the total number of node features associated with vertex i, and σ(·) is the nonlinear activation function;
[0027] S124. Aggregate new node features along the time axis to obtain SFP feature V1;
[0028]
[0029] In the formula, MaxPoeling T (·) represents max pooling. To enhance the historical trajectory after node features, This refers to the enhanced node features located in frame tK of the historical trajectory. For enhanced node features located in frame t-T+1 of the historical trajectory, This refers to the enhanced node features located in frame t-1 of the historical trajectory.
[0030] Further: In S1, the expression for obtaining SFC feature V3 is:
[0031]
[0032] In the formula, To combine the features of all enhanced target nodes in frame t, For the first enhanced node feature in frame t, For the second enhanced node feature in frame t, This is the enhanced node feature of the h-th frame t.
[0033] Further: S2 includes the following sub-steps:
[0034] S21. Input the current search image into the shared convolutional network to obtain the instance embedding of the current search image;
[0035] S22. Embed the instances of the current search image and input them sequentially into the convolutional layer and the max pooling layer to obtain global features;
[0036] S23. Input the global features into the deconvolution layer to obtain the context features.
[0037] Further: In S22, the size of the global feature is 256×1, the convolutional layer is set with 256 filters, the kernel size is 3×3, the stride is 1, and the size of the max pooling layer is 484.
[0038] Further: S3 includes the following sub-steps:
[0039] S31. The SFP features, SFC features, and context features are fused by element-wise addition to obtain the fused features;
[0040] S32. Generate an adaptive map based on the fusion features;
[0041] S33. Input the adaptive graph and SFP features into the context feature network to obtain the adaptive features;
[0042] S34. Based on the adaptive features and the instance embedding of the current search image, the correlation is calculated using the Xcorr function to obtain the predicted location of each target in the current search image.
[0043] Furthermore: In S32, the adjacency matrix of the adaptive graph The expression is:
[0044]
[0045] In the formula, N h V represents the number of targets in the current frame. x,n For fusion feature V x The nth column vector, V x,m For fusion feature V xThe m-th column vector, g(·) and h(·) are both 1×1 convolutional layers with 256 filters.
[0046] The beneficial effects of this invention are as follows:
[0047] (1) This invention provides a multi-target tracking method based on interactive spatiotemporal features of graph neural networks. By adding spatiotemporal graph attention network and context feature network to the network architecture, the spatiotemporal appearance modeling of the target, the interaction between targets, and context-guided adaptive learning can be combined to achieve robust target localization and tracking. Furthermore, based on the characteristic that different targets have different degrees of correlation, the attention mechanism is used to make the model focus on important nodes and edges when processing graph data, which effectively improves the tracking accuracy.
[0048] (2) This invention simultaneously considers the interactions between multiple targets within the same frame and the spatiotemporal feature model based on multiple targets across frames. This makes it more robust for target appearance modeling. It considers learning contextual feature information in the current frame by utilizing the relationships between targets. In context-guided feature adaptation, the current search image provides useful foreground / background information, which helps in feature-adaptive graphical learning and achieves scale adaptation of the tracked target. Robust target localization can be achieved using adaptive features. Furthermore, by combining these two spatiotemporal graphs, interactive spatiotemporal features can be comprehensively learned to represent the target's appearance.
[0049] (3) In the case of different nodes in the graph having no weight, the present invention aggregates the spatiotemporal features of the interaction of the target by a graph neural network with an attention mechanism, so that the nodes in the graph can make full use of the topological information of the graph and obtain more robust tracking results. Attached Figure Description
[0050] Figure 1 This is a flowchart of a multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to the present invention.
[0051] Figure 2 The network principle of this invention Figure 1 .
[0052] Figure 3 The network principle of this invention Figure 2 . Detailed Implementation
[0053] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0054] like Figure 1 As shown, in one embodiment of the present invention, a multi-target tracking method based on a graph neural network with interactive spatiotemporal features includes the following steps:
[0055] S1. Obtain historical sample images and the current search image, and input the historical sample images and the current search image into the feature extraction network to obtain SFP features and SFC features; wherein, the feature extraction network includes an interconnected shared convolutional network and a spatiotemporal graph attention network (SF-GAT);
[0056] S2. Input the current search image into a shared convolutional network, a max pooling layer, and a deconvolutional layer in sequence to obtain contextual features;
[0057] S3. Adaptive features are generated based on SFP features, SFC features and context features through the context feature network to obtain the predicted position of each target in the current search image, thus completing multi-target tracking.
[0058] The network architecture of this invention is as follows: Figures 2-3 As shown, this architecture integrates a spatiotemporal graph attention network and a context feature network into a Siamese network framework for target appearance modeling. It uses a spatiotemporal graph attention network to model the structuring of historical sample images and to model the target sample representation that fully considers the interactions between targets. In addition, by setting up a context graph convolutional network, it utilizes the context learning adaptive features of the current framework for target localization, thereby achieving high-performance visual tracking.
[0059] In step S1, the feature extraction network processes historical sample images and the current search image using the same method. Specifically, the feature extraction network processes historical sample images as follows:
[0060] S11. Input the historical sample images into the shared convolutional network to obtain the feature set;
[0061] S12. Construct an undirected graph based on the feature set, input the undirected graph into a spatiotemporal graph attention network to obtain a refined feature set, and aggregate the refined feature set to obtain SFP features.
[0062] like Figures 2-3 As shown, historical sample images are input into a shared convolutional network, and the generated feature set is specifically a cross-frame sample embedding. In order to perform spatiotemporal modeling of the target object, this invention uses T consecutive frames, where each frame contains N... Z An undirected graph is constructed on the sample embedding sequence of each node.
[0063] In S11, the historical sample images include several consecutive video frames;
[0064] The specific method for obtaining the feature set is as follows:
[0065] Based on several consecutive video frames and the target detection features in each frame, a feature set is obtained by embedding features of the input target object into different target regions in all video frames through a shared convolutional network.
[0066] In S12, the undirected graph includes a node set and an edge set. The node set is composed of all target parts in the embedding sequence of the feature set. The edge set includes first and second types of edges. The first type of edge is a spatial edge, which represents the intra-sample connection in each frame. The method to obtain the second type of edge is to connect the parts in consecutive frames that are at the same position as the temporal edge.
[0067] In this embodiment, considering that the edge set includes the first and second types of edges, the present invention uses an attention mechanism to limit each node to selecting only the N with the largest weight. Z The presence of a few nodes makes the constructed spatiotemporal graph sparse, reducing the computational cost of graph convolution operations.
[0068] In S12, the refined feature set includes vertex features and neighbor features. The specific method for obtaining SFP features is as follows:
[0069] S121. Calculate vertex features Characteristics of neighbors similarity coefficient e ij ;
[0070]
[0071] In the formula, W represents the shared parameters, || represents the feature concatenation operation, and a(·) represents the feature mapping operation. Let j be the set of node features adjacent to vertex i, where j is the index of the node feature;
[0072] In S121, the linear mapping of the shared parameter W increases the dimension of the vertex features, || concatenates the transformed features of the vertex features, and a(·) maps the concatenated high-dimensional features to a real number.
[0073] S122. Normalize the attention coefficients based on the similarity coefficients to obtain the normalized attention coefficients α. ij ;
[0074]
[0075] In the formula, e ik Let be the attention coefficient, LeaKerReLU(·) be the activation function, and k be the index of the neighboring node of vertex i.
[0076] S123. The node features to be aggregated are weighted and summed according to the normalized attention coefficients, and then enhanced through a multi-head attention mechanism to obtain enhanced new node features.
[0077]
[0078] In the formula, || represents the connection operation, K is the total number of node features associated with vertex i, and σ(·) is the nonlinear activation function; node features The overall formula represents the enhancement of a node feature by utilizing K node features associated with vertex i.
[0079] S124. Aggregate new node features along the time axis to obtain SFP feature V1;
[0080]
[0081] In the formula, MaxPoeling T (·) represents max pooling. To enhance the historical trajectory after node features, This refers to the enhanced node features located in frame tK of the historical trajectory. For enhanced node features located in frame t-T+1 of the historical trajectory, This refers to the enhanced node features located in frame t-1 of the historical trajectory.
[0082] In this embodiment, in order to reduce the computational burden of subsequent layers, the present invention aggregates features along the time axis to obtain compact SFP features V1.
[0083] In S1, the expression for SFC feature V3 is obtained as follows:
[0084]
[0085] In the formula, To combine the features of all enhanced target nodes in frame t, For the first enhanced node feature in frame t, For the second enhanced node feature in frame t, This is the enhanced node feature of the h-th frame t.
[0086] Considering the relationship between targets in the same frame, this invention uses a shared convolutional network to extract a feature set representation from the current search image to obtain an undirected graph. This undirected graph uses the features of multiple targets in a frame as the initial features of multiple nodes in the graph, and uses the interaction features between them as the edge connection representation of the interaction feature graph. Using the method of generating SFP features, the spatiotemporal graph attention network is used to obtain SFC feature V3.
[0087] S2 includes the following steps:
[0088] S21. Input the current search image into the shared convolutional network to obtain the instance embedding of the current search image;
[0089] S22. Embed the instances of the current search image and input them sequentially into the convolutional layer and the max pooling layer to obtain global features;
[0090] S23. Input the global features into the deconvolution layer to obtain the context features.
[0091] like Figures 2-3 As shown, the network framework of this invention not only models the spatiotemporal structure between target samples, but also incorporates the contextual features of the current search image to guide adaptive feature learning. To fully utilize contextual information, a graph learning model is integrated into the framework, which generates an adaptive graph to guide the contextual feature network.
[0092] In S22, the size of the global feature is 256×1, the convolutional layer has 256 filters, the kernel size is 3×3, the stride is 1, and the size of the max pooling layer is 484.
[0093] S3 includes the following steps:
[0094] S31. The SFP features, SFC features, and context features are fused by element-wise addition to obtain the fused features;
[0095] S32. Generate an adaptive map based on the fusion features;
[0096] S33. Input the adaptive graph and SFP features into the Context Feature Network (SF-GCN) to obtain the adaptive features;
[0097] S34. Based on the adaptive features and the instance embedding of the current search image, the correlation is calculated using the Xcorr function to obtain the predicted location of each target in the current search image.
[0098] In step S31, the fusion feature V is obtained. x The specific expression is:
[0099]
[0100] In the formula, The generated fusion feature V is based on contextual features. x It takes into account both the spatiotemporal characteristics of the target object and the interaction characteristics between targets, as well as the context information of the current frame.
[0101] In step S32, in order to perform graph learning to achieve robust feature adaptation, this invention uses fused features V.x Generate an adaptive graph, and the adjacency matrix of the adaptive graph. The expression is:
[0102]
[0103] In the formula, N h V represents the number of targets in the current frame. x,n For fusion feature V x The nth column vector, V x,m For fusion feature V x The m-th column vector, g(·) and h(·) are both 1×1 convolutional layers with 256 filters.
[0104] The beneficial effects of this invention are as follows: This invention provides a multi-target tracking method based on interactive spatiotemporal features of graph neural networks. By incorporating spatiotemporal graph attention networks and context feature networks into the network architecture, it can combine spatiotemporal appearance modeling of targets, interactions between targets, and context-guided adaptive learning to achieve robust target localization and tracking. Furthermore, based on the characteristic that different targets have different degrees of correlation, it utilizes an attention mechanism to make the model focus on important nodes and edges when processing graph data, effectively improving the tracking accuracy.
[0105] This invention simultaneously considers the interactions between multiple targets within the same frame and the spatiotemporal feature model based on multiple targets across frames. This makes it more robust for target appearance modeling. It considers learning contextual feature information in the current frame by utilizing the relationships between targets. In context-guided feature adaptation, the current search image provides useful foreground / background information, which helps in feature-adaptive graph learning and achieves scale adaptation of the tracked target. Using adaptive features, robust target localization can be achieved. Furthermore, by combining these two spatiotemporal graphs, interactive spatiotemporal features can be comprehensively learned to represent the target's appearance.
[0106] This invention addresses the issue of unweighted nodes in a graph by aggregating the spatiotemporal features of target interactions using a graph neural network with an attention mechanism. This allows nodes in the graph to fully utilize the graph's topological information, resulting in more robust tracking results.
[0107] In the description of this invention, it should be understood that the terms "center," "thickness," "upper," "lower," "horizontal," "top," "bottom," "inner," "outer," and "radial," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying the relative importance or the number of technical features implicitly specified. Therefore, a feature defined by "first," "second," and "third" may explicitly or implicitly include one or more of that feature.
Claims
1. A multi-target tracking method based on graph neural networks with interactive spatiotemporal features, characterized in that, Includes the following steps: S1. Obtain historical sample images and the current search image, and input the historical sample images and the current search image into the feature extraction network respectively to obtain SFP features and SFC features; wherein, the feature extraction network includes an interconnected shared convolutional network and a spatiotemporal graph attention network; S2. Input the current search image into a shared convolutional network, a max pooling layer, and a deconvolutional layer in sequence to obtain contextual features; S3. Adaptive features are generated based on SFP features, SFC features and context features through the context feature network to obtain the predicted position of each target in the current search image, thus completing multi-target tracking; In step S1, the feature extraction network processes historical sample images and the current search image using the same method. Specifically, the feature extraction network processes historical sample images as follows: S11. Input the historical sample images into the shared convolutional network to obtain the feature set; S12. Construct an undirected graph based on the feature set, input the undirected graph into a spatiotemporal graph attention network to obtain a refined feature set, and aggregate the refined feature set to obtain SFP features. In S12, the undirected graph includes a node set and an edge set. The node set is composed of all target parts in the embedding sequence of the feature set. The edge set includes first and second type edges. The first type of edge is a spatial edge, which represents the intra-sample connection in each frame. The method to obtain the second type of edge is to connect the parts in consecutive frames that are at the same position as the temporal edge. In S12, the refined feature set includes vertex features and neighbor features. The specific method for obtaining SFP features is as follows: S121. Calculate vertex features Characteristics of neighbors similarity coefficient ; In the formula, W To share parameters, For feature splicing operations, a (·) represents the feature mapping operation. For the vertex i The set of features of adjacent nodes. j The index of the node feature; S122. Normalize the attention coefficients based on the similarity coefficients to obtain the normalized attention coefficients. ; In the formula, Attention coefficient For activation function, k For the vertex i The sequence number of the adjacent neighbor node; S123. The node features to be aggregated are weighted and summed according to the normalized attention coefficients, and then enhanced through a multi-head attention mechanism to obtain enhanced new node features. ; In the formula, || represents the join operation. K For the vertex i The total number of associated node features, It is a non-linear activation function; S124. Aggregate new node features along the time axis to obtain SFP features. ; In the formula, For max pooling, To enhance the historical trajectory after node features, Located in the historical trajectory t - K Enhanced node features of frames, Located in the historical trajectory t - T Enhanced node features of +1 frame Located in the historical trajectory t Enhanced node features for frame -1.
2. The multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 1, characterized in that, In S11, the historical sample images include several consecutive video frames; The specific method for obtaining the feature set is as follows: Based on several consecutive video frames and the target detection features in each frame, a feature set is obtained by embedding features of the input target object into different target regions in all video frames through a shared convolutional network.
3. The multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 1, characterized in that, In S1, the SFC feature is obtained. The expression is: In the formula, To combine the features of all enhanced target nodes in frame t, For the first enhanced node feature in frame t, For the second enhanced node feature in frame t, For the t-th frame h An enhanced node feature.
4. The multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 1, characterized in that, S2 includes the following steps: S21. Input the current search image into the shared convolutional network to obtain the instance embedding of the current search image; S22. Embed the instances of the current search image and input them sequentially into the convolutional layer and the max pooling layer to obtain global features; S23. Input the global features into the deconvolution layer to obtain the context features.
5. The multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 4, characterized in that, In S22, the size of the global feature is 256×1, the convolutional layer has 256 filters, the kernel size is 3×3, the stride is 1, and the size of the max pooling layer is 484.
6. The multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 5, characterized in that, S3 includes the following steps: S31. The SFP features, SFC features, and context features are fused by element-wise addition to obtain the fused features; S32. Generate an adaptive map based on the fusion features; S33. Input the adaptive graph and SFP features into the context feature network to obtain the adaptive features; S34. Based on the adaptive features and the instance embedding of the current search image, the correlation is calculated using the Xcorr function to obtain the predicted location of each target in the current search image.
7. A multi-target tracking method based on interactive spatiotemporal features of a graph neural network according to claim 6, characterized in that, In S32, the adjacency matrix of the adaptive graph The expression is: In the formula, This indicates the number of targets in the current frame. For fusion features The n Column vector, For fusion features The m The column vectors g(·) and h(·) are both 1 × 1 convolutional layers with 256 filters.
Citation Information
Patent Citations
Target tracking method and system based on context self-attention learning deep network
CN116109678A
Target tracking method based on space-time interaction attention mechanism
CN116563355A