Interrupted target track continuation association method based on spatiotemporal attention metric network
By introducing recurrent attention and spatial attention mechanisms into the spatiotemporal metric network, the weights of key attributes and nodes of target tracks are amplified, which solves the problem of continuous association of interrupted target tracks under dense target distribution and high track similarity, and achieves higher association accuracy and reliability.
Patent Information
- Application Number
- CN202310791168.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-06-29
AI Technical Summary
Existing multi-sensor target track association methods have difficulty in effectively extracting the discriminative features of target tracks in scenarios where targets are densely distributed and track similarity is high, resulting in poor performance in resuming the association of interrupted target tracks.
A temporal information extraction module with a recurrent attention mechanism and a spatial information extraction module with a spatial attention mechanism are adopted, combined with a symmetric loss function and an adversarial loss function. The key attributes and node weights of the target track are amplified through the spatiotemporal metric network, the information loss in the network learning process is reduced, and the accurate association of the interrupted target track is achieved.
In scenarios where targets are densely distributed and track similarity is high, the accuracy and reliability of track continuation of interrupted targets are improved, and the information loss during network learning is reduced.
Smart Images

Figure CN117033915B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-sensor target track association and fusion, and in particular relates to a method for reconnecting an interrupted target track. Background Art
[0002] Multi-sensor target track association is a technique for determining the origin of multi-sensor target tracks and is a prerequisite and key to multi-sensor target track fusion. Due to limitations in sensor detection performance, the target's high motion, and environmental occlusion, target tracks can be interrupted across regions. This results in discontinuities in the temporal and spatial distribution of the target track, hindering subsequent multi-sensor target track fusion and tracking. Track continuation of interrupted targets involves correlating the old track before and after the interruption of the same target. Existing multi-sensor target track association methods are primarily categorized as track prediction and data-driven methods. Track prediction methods, based on certain target prior information, predict the old track before the interruption and match it with the new track of the target to be associated after the interruption to achieve continuous association of the interrupted target track. This type of method relies excessively on prior information about the scene and target, making it unsuitable for scenarios with uncertain and complex target variability. Data-driven methods, based on offline data, use neural networks to learn deep feature representations of the old and new tracks before and after the interruption, transforming the problem of continuous association of interrupted target tracks into a problem of matching corresponding feature representations. Compared with the trajectory prediction method, the data-driven method only requires a sufficient amount of target trajectory data for learning network parameters and has better correlation performance.
[0003] Metric learning network-based track continuation is a data-driven method for track continuation of interrupted targets. Specifically, it extracts key features from the old and new tracks before and after the target interruption using a deep neural network, representing them as a unified feature vector. The network is trained using the principle of decreasing the distance between feature vectors of the same target and increasing the distance between feature vectors of different targets, so that the extracted key features can fully represent the identity characteristics of the original corresponding tracks. This network is called a metric learning network. Based on the trained network model, binary classification can be used to continuation of a pair of target tracks. Typically, metric learning networks used to extract key target features consist of convolutional neural networks and recurrent neural networks, which can fully represent the temporal and spatial information of target tracks. However, when targets are densely distributed and track similarity is high, basic spatiotemporal metric networks struggle to extract discriminative features of target tracks, resulting in errors and difficulties in target track continuation. Therefore, the present invention adds two attention mechanisms with different functions into the spatiotemporal metric network, amplifies the weights of key attributes and nodes of target tracks during network learning, improves the ability to extract key features of target tracks, and can realize the continuation of interrupted target tracks in scenarios with dense target distribution and high track similarity.
[0004] The existing technical solutions for connecting interrupted target tracks based on metric learning networks are as follows:
[0005] (1) First, the temporal information features of a pair of new and old tracks are extracted, and the original target track is converted into a spatial feature map F representing the feature relationship between the target track nodes through a single-layer LSTM structure.
[0006] (2) The spatial feature map F is subjected to a multi-scale residual convolution kernel to extract spatial features and mapped into a feature vector of size N.
[0007] (3) For all training samples of new and old track pairs, repeat steps 1 and 2 according to the contrast loss function and whether the new and old tracks originate from the same target.
[0008] (4) For all the new and old track pair test samples, they are input into the network in sequence, the target track feature vector association threshold M is set, the binary classification results are output, and the target track connection association is completed.
[0009] The key to reconnecting interrupted target tracks lies in the target track temporal information feature extraction module in step (1) and the target track spatial information feature extraction module in step (2). Xiong et al. replaced the spatiotemporal metric network with a graph neural network, which improved the ability to extract the topological relationship of the target track. However, the target track is generally a simple two-dimensional or three-dimensional sequence, and its topological relationship is a simple linear connection relationship, which does not significantly improve the ability to extract the key features of the target track.
[0010] When targets are densely distributed and the similarity of target tracks is high, existing methods can only extract the overall characteristics of target tracks through basic LSTM modules and residual convolution modules, resulting in similar features extracted from different target tracks in the spatiotemporal metric network. In addition, the spatial information extraction module is established on the basis of the temporal information extraction module, resulting in information loss in the output of the temporal feature extraction module, and the performance of the connection between the old and new tracks before and after the target interruption is poor.
[0011] In summary, among the existing methods for reconnecting interrupted target tracks based on metric learning networks:
[0012] (1) Only focusing on the overall macroscopic characteristics of the target track, lacking a mechanism to focus on its local microscopic characteristics;
[0013] (2) Information loss occurs in the network intermediate modules. Summary of the Invention
[0014] In order to overcome the shortcomings of the existing technology, the present invention provides a method for connecting and associating interrupted target tracks based on a spatiotemporal attention metric network. The method designs a target track temporal information extraction module with a recurrent attention mechanism and a target track spatial information extraction module with a spatial attention mechanism, amplifies the weights of key attributes and nodes of the target track during the network learning process, and ensures the symmetry of feature weights between nodes through a symmetric loss function. In addition, in order to reduce information loss during the network learning process, the present invention also splices the output of the temporal information extraction module with the output of the spatial information extraction module to form a final feature vector. The method of the present invention reduces information loss during the network learning process and achieves excellent interrupted target track connection and association performance.
[0015] The technical solution adopted by the present invention to solve the technical problem includes the following steps:
[0016] Step 1: For the target complete track set T = {T1,…,T n} is truncated, the truncation interval f represents the interruption length of the track, and the truncation length l represents the length of the track after truncation, thereby constructing the target track set T before interruption old ={T old_1 ,…,T old_n} and the post-interruption track set T new ={T new_1 ,…,T new_n};
[0017] Step 2: Select a set of tracks from each of the two track sets {T old ,T new}, repeat several times to construct the target track pair training set T train ={T old ,T new} N and the test set T train ={T old ,T new} n-N , and perform MIN-MAX processing on the data of the two parts of the training set;
[0018] Step 3: Align the target track in the training set to T train ={T old ,T new} N Input the spatiotemporal metric network with attention mechanism to learn the discontinuous spatiotemporal features that can continue the interrupted track and convert the learned features into a feature vector of size N. The specific steps are as follows:
[0019] Step 3-1: The target track pair passes through a recurrent attention mechanism module, which consists of two layers of ReLU and one layer of softmax, to generate weights for each attribute of the corresponding input track pair. The attribute weights are weighted summed with the input track pair to ensure that the weights of attributes that are more important to the target are amplified during the network learning process.
[0020] Step 3-2: The new target track pair obtained in step 3-1 is converted into a spatial feature map between target track nodes through a time series information extraction module with a three-layer LSTM. Each position in the spatial feature map represents the feature relationship between the corresponding position nodes and has symmetry.
[0021] Step 3-3: The spatial feature map obtained in step 3-2 is passed through the spatial attention mechanism module, which consists of a 3×3 convolution kernel and a sigmoid layer. The module generates symmetrical spatial attention weights to indicate the importance of the feature relationship between nodes. At the same time, the spatial attention weights are weighted and summed with the spatial feature map.
[0022] Step 3-4: The spatial feature map after the weighting of the dual attention mechanism is passed through 1×1 and 3×3 convolution kernels for multi-scale spatial feature extraction;
[0023] Step 3-5: The multi-scale extracted features are passed through three layers of 3×3 residual convolution kernels to further extract the features, and then converted into a feature vector of size N through a fully connected layer;
[0024] Step 4: Use the adversarial loss function to constrain the goal and direction of network learning. Define Q1 and Q2 as the feature vectors output by steps 3-5 for the two tracks, D(Q1, Q2) represents the Euclidean distance between them, and l is the identity label of the interrupted track pair. When a pair of interrupted tracks belong to the same complete track, l = 1, otherwise l = 0; a ij and w ij They represent the corresponding position values of the spatial feature map in step 3-2 and the spatial attention weight matrix in step 3-3 respectively; the adversarial loss function is set as follows:
[0025]
[0026] Under the constraints of the loss function, the network parameter training is repeated repeatedly;
[0027] Step 5: In the trained network, input the target track pair TE1 and TE2 in the test set, and determine whether the corresponding track pairs are associated based on the distance between the feature vectors Q1 and Q2 corresponding to the two tracks and the target association threshold M.
[0028] Furthermore, the method for determining whether the corresponding track pairs are associated in step 5 is as follows:
[0029]
[0030] The beneficial effects of the present invention are as follows:
[0031] The present invention basically does not rely on scene and target prior information. In scenarios with dense target distribution and high track similarity, it considers the different effects of different attributes and nodes in the target track on the continuity association, injects recurrent attention mechanism and spatial attention mechanism into the spatiotemporal metric network, reduces the information loss in the network learning process, and achieves excellent interrupted target track continuity association performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is the overall framework diagram of the present invention.
[0033] Figure 2 This is the temporal information extraction method of the recurrent attention mechanism of the present invention.
[0034] Figure 3 This is the spatial information extraction method of the spatial attention mechanism of the present invention. DETAILED DESCRIPTION
[0035] The present invention will be further described below with reference to the accompanying drawings and examples.
[0036] The purpose of the present invention is to propose a method for continuing the tracks of interrupted targets based on a spatiotemporal metric network with an attention mechanism under the constraints of scenarios with dense target distribution and high track similarity. Two attention mechanisms with different functions are added to the spatiotemporal metric network, which amplifies the weights of key attributes and nodes of target tracks during network learning, improves the ability to extract key features of target tracks, and realizes the continuation of interrupted target tracks in scenarios with dense target distribution and high track similarity.
[0037] like Figure 1 As shown in FIG, a method for reconnecting interrupted target tracks based on a spatiotemporal attention metric network includes the following steps:
[0038] Step 1: For the target complete track set T = {T1,…,T n} is truncated, the truncation interval f represents the interruption length of the track, and the truncation length l represents the length of the track after truncation, thereby constructing the target track set T before interruption old ={T old_1 ,…,T old_n} and the post-interruption track set T new ={T new_1 ,…,T new_n};
[0039] Step 2: Select a set of tracks from each of the two track sets {T old ,T new}, repeat several times to construct the target track pair training set T train ={T old ,T new} N and the test set T train ={T old ,T new} n-N , and perform MIN-MAX processing on the data of the two parts of the training set;
[0040] Step 3: Align the target track in the training set to T train ={T old ,T new} N The input network learns the discontinuous spatiotemporal features that can continue the interrupted track through the spatiotemporal metric network with attention mechanism, and converts the learned features into a feature vector of size M. The specific steps are as follows:
[0041] Step 301: The target track pair passes through a recurrent attention mechanism module, which consists of two layers of ReLU and one layer of softmax, to generate weights for each attribute of the corresponding input track pair. The attribute weights are weighted and summed with the input track pair to ensure that the weights of attributes that are more important to the target are amplified during the network learning process; Figure 2 As shown;
[0042] Step 302: The new target track pair obtained in step 301 is converted into a spatial feature map between target track nodes through a time series information extraction module with a three-layer LSTM. Each position in the spatial feature map represents the feature relationship between the corresponding position nodes and has symmetry.
[0043] Step 303: The spatial feature map obtained in step 302 is passed through a spatial attention mechanism module, which consists of a 3×3 convolution kernel and a sigmoid layer. The module generates symmetrical spatial attention weights to indicate the importance of feature relationships between nodes. At the same time, the spatial attention weights are weighted and summed with the spatial feature map to ensure that the weights of nodes that are more important to the target are amplified during the network learning process. Figure 3 shown.
[0044] Step 304: The spatial feature map after the weighting of the dual attention mechanism is subjected to 1×1 and 3×3 convolution kernels respectively for multi-scale spatial feature extraction to ensure that features of different granularities are fully extracted.
[0045] Step 305: The multi-scale extracted features are passed through three layers of 3×3 residual convolution kernels to further extract the features, and then converted into a feature vector of size M through a fully connected layer;
[0046] Step 4: Use the adversarial loss function to constrain the target and direction of network learning. First, ensure that the distance between the feature vectors corresponding to the same target track is reduced, and the distance between the feature vectors corresponding to different target tracks is increased; secondly, the symmetry of the spatial feature map generated in step 302 must be satisfied; finally, the symmetry of the spatial attention weights in step 303 must be satisfied. The adversarial loss function is set as follows:
[0047]
[0048] Under the constraints of the loss function, the network parameter training is repeated repeatedly;
[0049] Step 5: In the trained network, input the target track pair in the test set and determine whether the corresponding track pair is associated based on the target association threshold:
[0050] Specific embodiment:
[0052] The method of the present invention is described by taking the KITTI_tracking_label dataset as an example:
[0053] Step 1: Using the interruption interval d and the track length L, the complete target track is truncated to construct the old track set T1 before the target interruption and the new track set T2 after the interruption;
[0054] Step 2: Select a set of tracks from track set T1 and track set T2 respectively, repeat several times, and construct the target track pair training set TR and test set TE;
[0055] Step 3: Input a pair of target track pairs TR1 and TR2 from the training set TR into the spatiotemporal metric network with an attention mechanism to extract the spatiotemporal feature information of the target track and convert it into feature vectors Q1 and Q2 of size N. The detailed steps are as follows:
[0056] 3-1. For a target track pair TR1 and TR2 with an attribute dimension of 2 or 3 and a length of L, a cyclic attention mechanism module is used to generate 2-dimensional or 3-dimensional attribute weights W1 and W2. The attribute weights are weighted and summed with the original target track pair to obtain new target tracks TR1' and TR2' with amplified attribute weights. The calculation process is as follows:
[0057] h1=ReLU(W1·TR+b1)
[0058] h2=ReLU(W2·h1+b2)
[0059] W x =softmax(h2)
[0060] TR'=W x TR
[0061] 3-2. Convert TR1' and TR2' obtained in step 3-1 into a spatial feature map F of size L×L through a three-layer LSTM structure. This map represents the corresponding feature relationship between the L nodes of the target track and has symmetry. The symmetry of F is ensured by using a symmetric loss function. The symmetric loss function of F is as follows:
[0062]
[0063] 3-3. The spatial feature map F of size L×L is passed through a 3×3 convolution kernel and a sigmoid layer to generate the spatial attention weight W, which represents the importance of the feature relationship between nodes. It is symmetrical and is maintained by a symmetric loss function. At the same time, W and F are weighted and summed to obtain a new spatial feature map F' that amplifies the node weights. The symmetric loss function of W is as follows:
[0064]
[0065] 3-4. The spatial feature map F' of the amplified node weight undergoes the first round of multi-scale spatial feature extraction using 1×1 and 3×3 convolution kernels, then undergoes the second round of spatial feature extraction using the residual 3×3 convolution kernel, and finally is converted into a feature vector of size N through a fully connected layer.
[0066] Step 4: Based on whether TR1 and TR2 originate from the same target, the adversarial loss function is used to ensure that the distances Q1 and Q2 corresponding to the same target track pair are reduced, while the distances Q1 and Q2 corresponding to different target track pairs are increased. The network parameter training is repeated repeatedly. The network loss function includes the adversarial loss function and the two-part symmetric loss function:
[0067]
[0068] Step 5: Input a pair of target track pairs TE1 and TE2 from the test set TE into the trained network. Based on the target association threshold M and the distance between the network outputs Q1 and Q2, perform binary classification to determine whether the corresponding track pairs are associated:
[0069]
Claims
1. A method for reconnecting interrupted target tracks based on spatiotemporal attention metric network, characterized in that: The steps include: Step 1: For the target complete track set T = {T1,…,T n } is truncated, the truncation interval f represents the interruption length of the track, and the truncation length l represents the length of the track after truncation, thereby constructing the target track set T before interruption old ={T old_1 ,…,T old_n } and the post-interruption track set T new ={T new_1 ,…,T new_n }; Step 2: Select a set of tracks from each of the two track sets {T old ,T new }, repeat several times to construct the target track pair training set T train ={T old ,T new } N and the test set T train ={T old ,T new } n-N , and perform MIN-MAX processing on the data of the two parts of the training set; Step 3: Align the target track in the training set to T train ={T old ,T new } N Input the spatiotemporal metric network with attention mechanism to learn the discontinuous spatiotemporal features that can continue the interrupted track and convert the learned features into a feature vector of size N. The specific steps are as follows: Step 3-1: The target track pair passes through a recurrent attention mechanism module, which consists of two layers of ReLU and one layer of softmax to generate weights for each attribute of the corresponding input track pair; The attribute weights are weighted and summed with the input track pairs to ensure that the attribute weights that are more important to the target are amplified during the network learning process; Step 3-2: The new target track pair obtained in step 3-1 is converted into a spatial feature map between target track nodes through a time series information extraction module with a three-layer LSTM. Each position in the spatial feature map represents the feature relationship between the corresponding position nodes and has symmetry. Step 3-3: The spatial feature map obtained in step 3-2 is passed through the spatial attention mechanism module, which consists of a 3×3 convolution kernel and a sigmoid layer. The module generates symmetrical spatial attention weights to indicate the importance of the feature relationship between nodes. At the same time, the spatial attention weights are weighted and summed with the spatial feature map. Step 3-4: The spatial feature map after the weighting of the dual attention mechanism is passed through 1×1 and 3×3 convolution kernels for multi-scale spatial feature extraction; Step 3-5: The multi-scale extracted features are passed through three layers of 3×3 residual convolution kernels to further extract the features, and then converted into a feature vector of size N through a fully connected layer; Step 4: Use the adversarial loss function to constrain the goal and direction of network learning. Define Q1 and Q2 as the feature vectors output by steps 3-5 for the two tracks, D(Q1, Q2) represents the Euclidean distance between them, and l is the identity label of the interrupted track pair. When a pair of interrupted tracks belong to the same complete track, l = 1, otherwise l = 0; a ij and w ij They represent the corresponding position values of the spatial feature map in step 3-2 and the spatial attention weight matrix in step 3-3 respectively; the adversarial loss function is set as follows: Under the constraints of the loss function, the network parameter training is repeated repeatedly; Step 5: In the trained network, input the target track pair TE1 and TE2 in the test set, and determine whether the corresponding track pairs are associated based on the distance between the feature vectors Q1 and Q2 corresponding to the two tracks and the target association threshold M.
2. The method for connecting interrupted target tracks based on a spatiotemporal attention metric network according to claim 1, characterized in that: The method for determining whether the corresponding track pairs are associated in step 5 is as follows:
Citation Information
Patent Citations
Interrupted track continuing association deep learning method
CN111898746A
High-resolution remote sensing image change detection method based on global relation reasoning
CN114913434A