A Vehicle Trajectory Tracking Method Based on Spatiotemporal Interaction Neural Network

By adopting spatiotemporal interaction neural networks and improved deep Hungarian networks in vehicle trajectory tracking, tracking errors caused by manual operation lag and sensor equipment limitations in traditional methods are solved, and efficient and accurate vehicle trajectory prediction and tracking are achieved.

CN114782489BActive Publication Date: 2025-06-20ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210387072.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-06-20
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The existing vehicle tracking methods have lagged processing time due to manual attention and speed problems, and the sensor itself has equipment limitations and defects, resulting in output of incorrect tracking results, especially in poor performance in vehicle driving at intersections.

Method used

The vehicle trajectory tracking method based on spatiotemporal interaction neural network is adopted, and the vehicle coordinates are encoded using long and short-term memory networks and graph attention networks to predict the coordinates of the vehicle's next moment, and the prediction network is trained through the improved deep Hungarian network to improve tracking accuracy in an end-to-end manner.

Benefits of technology

It realizes that the accuracy and efficiency of vehicle trajectory tracking is improved based on time and space factors, and the lag of manual operation and the error rate caused by sensor equipment limitation are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782489B_ABST
    Figure CN114782489B_ABST
Patent Text Reader

Abstract

A vehicle trajectory tracking method based on a spatio-temporal interaction neural network, comprising: First, preprocess sensor data, convert relative coordinates into geographic coordinates, and perform clustering to obtain average geographic coordinate points; Secondly, match vehicle observations with the tracker trajectory and update the tracker trajectory record; Then, use a prediction model based on spatio-temporal interaction to perform a single-step prediction on the tracker's trajectory record and wait for the next round of vehicle observation matching; In the training phase, an improved deep Hungarian network is also used to calculate the current prediction tracking accuracy to adjust the parameters of the prediction model. The present invention can jointly process vehicle coordinates from multiple sensors; predict vehicle positions considering both vehicle time factors and space factors; and realize end-to-end training of the prediction model with tracking accuracy as the index by means of an improved deep Hungarian network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of trajectory tracking, and includes a vehicle trajectory tracking method considering time factors and space factors. Background Art

[0002] In the field of intelligent transportation, traditional methods usually display road conditions in the form of videos, and a large amount of manual operations are required to monitor and track targets. Due to the attention and speed problems of manual operations, this method shows a lag in processing time. With the integration of digital twin technology and the transportation field, a traffic simulation platform based on this technology has been developed. In order to achieve accurate simulation, the traffic simulation platform based on digital twin technology needs to obtain high-quality vehicle trajectory data. Currently, the traffic department usually obtains the road vehicle status mainly by non-intrusive methods, such as processing surveillance videos and sensor signals.

[0003] However, vehicles have variable motion patterns during driving, and there are also complex spatial and temporal interaction relationships between vehicles; on the other hand, some sensors have certain equipment limitations and defects. These two factors will both lead to incorrect tracking results output by these non-intrusive methods. And these errors are particularly prominent in vehicle driving at intersections. Existing research mainly focuses on improving the detector accuracy in the tracking method, and the research on trackers and data association algorithms is relatively backward. Recently, a study at the 2020 CVPR conference proposed an improvement to the data association algorithm using the Deep Hungarian Network (DHN).

[0004] On this basis, the present invention proposes a vehicle trajectory tracking method based on a spatio-temporal interaction neural network, which uses a long short-term memory network and a graph attention network to encode vehicle coordinates to obtain the coordinate prediction of the vehicle at the next moment, and uses an improved deep Hungarian network to train the prediction network, so as to train the tracking accuracy of the prediction model in an end-to-end manner to achieve effective vehicle trajectory tracking. Summary of the Invention

[0005] The purpose of the present invention mainly aims at the limitations of the above vehicle trajectory tracking method, and proposes a vehicle trajectory tracking method based on a spatio-temporal interaction neural network. The aim is to use a prediction network considering time factors and space factors to predict the position of the vehicle at the next moment, and use an improved deep Hungarian network and tracking metrics to train the prediction network end-to-end, and finally achieve an effective vehicle trajectory tracking method.

[0006] A vehicle trajectory tracking method based on a spatio-temporal interaction neural network of the present invention includes the following steps:

[0007] S1. Preprocess the sensor data, convert the relative coordinates to geographic coordinates, and cluster to obtain the average geographic coordinate points;

[0008] S2. Match the vehicle observations with the tracker prediction values and update the tracker trajectory records;

[0009] S3. Use the prediction model to perform a single-step prediction on the tracking trajectory of the tracker and enter the next round of matching;

[0010] S4. Use the improved deep Hungarian network to perform end-to-end training on the prediction model with the tracking accuracy as the goal.

[0011] The specific steps of step S1 include:

[0012] S1.1. Denote the relative coordinates obtained by the sensor as where represents the i-th coordinate obtained by the j-th sensor; convert it to an approximate geographic coordinate

[0013]

[0014]

[0015] where and represent the longitude and latitude of the geographic coordinates obtained after the vehicle conversion; lng0 and lat0 represent the longitude and latitude of the geographic coordinates of the sensor as the coordinate origin; R represents the radius of the earth in meters;

[0016] S1.2. Cluster the geographic coordinates after conversion of all sensors on the same road section at a distance of half a meter using a distance-based clustering algorithm, and then take the in-cluster average value:

[0017]

[0018] where, represents the in-cluster average geographic coordinate of cluster c as the representative of this cluster; is the number of in-cluster geographic coordinates of cluster c; and represent the longitude and latitude of the i-th point in cluster c respectively;

[0019] S1.3. Cluster the geographic coordinates after conversion of all sensors on the road network at a distance of half a meter using a distance-based clustering algorithm, and use equation (3) to take the in-cluster average value; finally, obtain the processed data coordinate point set P t , representing all vehicle observations obtained at time t.

[0020] The specific steps of step S2 include:

[0021] S2.1. Calculate for P t and the predicted value of the tracker Calculate the coordinate distance between each pair:

[0022]

[0023]

[0024] where d i,j represents the predicted coordinate point of the i-th tracker at time t and the coordinate point of the j-th vehicle at the same time The distance between them; D is a matrix of size representing the number of trackers at time t, representing the number of vehicle observations at time t; D i,j represents the value of the i-th row and j-th column in the matrix. When d i,j ≤ 1 (meter), take d i,j , otherwise take Inf to represent infinity;

[0025] S2.2. Use the Hungarian algorithm to calculate D to obtain the minimum cost matching between the tracker and the vehicle observations;

[0026] S2.3. For the vehicle observations with matching objects, add them to the matching tracker as the trajectory record of the tracker; for the vehicle observations without matching objects, generate a new tracker for each observation and add it to the tracker; for the trackers without matching objects, add their predictions at the current moment to the trajectory record. If the continuous time of the un-matched trackers exceeds the threshold ε, delete the tracker;

[0027] S2.4. For each tracker, two contents will be stored, namely the trajectory record Tr i and the hidden state H i , where H i can be split into four parts, namely and In the initialization stage, Tr i has one observation, and the values of H i are all set to 0, that is The specific steps of step S3 include:

[0028] S3.1. For the i-th tracker, take out the vehicle coordinate points at the latest two moments from the trajectory record and calculate the relative displacement of their longitude and latitude:

[0029]

[0030] Among them, and respectively represent the vehicle coordinate points of the i-th tracker trajectory record at the (t - 1) moment;

[0031] S3.2. Concatenate the relative displacements of longitude and latitude, which is the overall displacement of the vehicle represented by the i-th tracker at the (t - 1) moment Take it as the input of the long short-term memory network:

[0032]

[0033]

[0034] Among them, ‖ represents the concatenation operation; L-LSTM is the first LSTM in the present invention, called local LSTM, which is used to encode the time embedding of a single vehicle; and respectively represent the hidden state and cell state of the i-th tracker corresponding to the L-LSTM network at the t moment; W l and b l respectively represent the weights and biases of L-LSTM, which are shared by all vehicles at the same moment;

[0035] S3.3. Then take as the input of the graph attention network to obtain the attention coefficients of the implicit spatial interaction between vehicles:

[0036]

[0037] Among them, is the attention coefficient of node j to node i at the t moment; · T represents matrix transpose; represents the set of neighbor nodes of node i on the graph; W ∈ R F′×F is the weight matrix shared by all nodes, which is used to perform a linear transformation on the features of each node, where F represents the dimension of, and F′ is the output dimension; a ∈ R 2F′ is the shared attention mechanism used to calculate the attention coefficients; a undergoes matrix operations and is then normalized through the activation function LeackyReLU and the Softmax method;

[0038] S3.4. Use the attention coefficients for calculation to obtain the spatial embedding of the vehicle represented by the i-th tracker

[0039]

[0040] Among them, denotes the concatenation of K spatial attention embeddings; σ(·) denotes the activation function, which is a non-linear method; is the normalized attention coefficient calculated by the k-th attention mechanism, W k is the weight matrix of the corresponding input linear transformation;

[0041] S3.5. Take as the input of the second long short-term network:

[0042]

[0043] where G-LSTM is the second LSTM in the present invention, called the global LSTM, and is used to encode the temporal embedding of each vehicle after being implicitly spatially influenced by other vehicles; and respectively represent the hidden state and cell state of the i-th tracker corresponding to the G-LSTM network at time t; W g and b g respectively represent the weights and biases of G-LSTM, which are shared by all vehicles at the same moment;

[0044] S3.6. Add the local encoding and global encoding of each vehicle to obtain the output embedding of the vehicle represented by each tracker At this time is equivalent to the residual; then take this input embedding as the input of the fully connected layer to obtain the final result:

[0045]

[0046] where FC(·) represents the fully connected layer; W f and b f are respectively the weights and biases of the fully connected layer;

[0047] S3.7. After obtaining the relative displacement prediction of each tracker at time t, add it to the vehicle coordinate point of the tracker at time t-1 to obtain the tracker prediction value at time t

[0048] S4 specifically includes:

[0049] S4.1. In the training phase, first calculate the D matrix using a method similar to that shown in step S2.1:

[0050]

[0051] Among them, α is a scaling factor, and the value of the calculation result needs to be scaled to the range of [0, 0.1]. This value needs to be selected according to the specific application scenario;

[0052] S4.2. Use D as the input of the deep Hungarian model to obtain the soft assignment matrix The size of this matrix is the same as the input, and the output values are between [0, 1];

[0053] S4.3. When the correct matches are known, the true positive matrix B of D can be obtained TP , and the true positive one-dimensional matrices P r and P c :

[0054]

[0055]

[0056]

[0057] Among them, represents the value of B TP at the i-th row and j-th column. If the prediction value of the i-th tracker and the observation value of the j-th vehicle are true positive matches, then this value is 1, otherwise it is 0;

[0058] S4.3. According to the method of the deep Hungarian model, for add a row (or column) of data δ as a threshold in the row (or column) dimension respectively, and then calculate Softmax in the row (or column) dimension to obtain C r (or C c ); The present invention further processes it to obtain improved false positives and false negatives

[0059]

[0060] S4.4:. According to the method of the deep Hungarian model, the representing ID switching errors can be calculated, and then these three errors are calculated as tracking metrics MOTA and MOTP as loss functions:

[0061]

[0062]

[0063] Among them, M represents the number of vehicle observation values; γ is a control coefficient used to adjust Importance; ⊙ represents the Hadamard product; ‖·‖1 represents the sum of all values within the matrix, and ‖·‖0 represents the number of all non - zero values within the matrix; λ is a control coefficient used to adjust the importance of dMOTP.

[0064] The working principle of the present invention is as follows: Use coordinate system transformation and clustering algorithms to unify multiple sensor data, then use LSTM and GAT models to capture the spatio - temporal relationships of vehicles to achieve prediction, and finally use an improved DHN model to train the predictor with the tracking accuracy as the goal.

[0065] The advantages of the present invention are as follows: Convert the data of multiple sensors into the same coordinate system for unified processing; On the basis of considering the time correlation of the vehicle itself, further consider the spatial correlation between vehicles to improve the prediction effect; And use an improved deep Hungarian network to train the prediction network to achieve the tracking accuracy of training the prediction model in an end - to - end manner. Brief Description of the Drawings

[0066] Figure 1 is the structural diagram of the method of the present invention;

[0067] Figure 2 is the overall flowchart of data processing, vehicle - tracker matching, tracker prediction, and prediction model training of the method of the present invention. Detailed Embodiments

[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the detailed embodiments of the present invention in detail.

[0069] The embodiment of the present invention provides a method for generating and visualizing a vehicle evacuation plan based on SUMO simulation. The system structure is as Figure 1 shown, and the path generation process is as Figure 2 shown. The method includes:

[0070] S1. Pre - process the sensor data, convert the relative coordinates into geographical coordinates, and cluster to obtain the average geographical coordinate points, including the following steps:

[0071] S1.1. Denote the relative coordinates obtained by the sensor as where represents the i - th coordinate obtained by the j - th sensor; Convert it into an approximate geographical coordinate The error of the finally obtained approximate geographical coordinate does not exceed 0.1 meter;

[0072]

[0073]

[0074] where and represent the longitude and latitude of the geographical coordinates obtained by converting the vehicle; lng0 and lat0 represent the longitude and latitude of the geographical coordinates of the sensor, serving as the coordinate origin; R represents the radius of the earth, in meters;

[0075] S1.2. Cluster the geographical coordinates after conversion of the sensors on all the same road segments at a distance of half a meter using a distance-based clustering algorithm, and then take the in-cluster average:

[0076]

[0077] where, represents the average of the geographical coordinates within the cluster c of the clustering cluster, serving as the representative of this clustering cluster; is the number of geographical coordinates within the cluster c of the clustering cluster; and respectively represent the longitude and latitude of the i-th point within the cluster c;

[0078] S1.3. Cluster the geographical coordinates after conversion of all the sensors on the road network at a distance of half a meter using a distance-based clustering algorithm, and use Equation (3) to take the in-cluster average; finally, obtain the processed data coordinate point set P t , representing all the vehicle observations obtained at time t.

[0079] S2. Match the vehicle observations with the tracker prediction values and update the tracker trajectory records, including the following steps:

[0080] S2.1. Calculate the coordinate distances pairwise between P t and the prediction values of the tracker :

[0081]

[0082]

[0083] where, d i,j represents the predicted coordinate point of the i-th tracker at time t and the coordinate point of the j-th vehicle at the same time ; D is a matrix of size , represents the number of trackers at time t, represents the number of vehicle observations at time t; D i,j represents the value of the i-th row and j-th column in the matrix, and when d i,j ≤1 (meter), take d i,j , otherwise take Inf to represent infinity;

[0084] S2.2. Calculate D using the Hungarian algorithm to obtain the minimum-cost matching between the tracker and vehicle observations;

[0085] S2.3. For vehicle observations with matching objects, add them to the matching tracker as the trajectory record of the tracker; for vehicle observations without matching objects, create a new tracker for each observation and add it to the tracker; for trackers without matching objects, add their predictions at the current moment to the trajectory record, and if the continuous time of a tracker without a match exceeds the threshold ε, delete the tracker;

[0086] S2.4. For each tracker, two contents are stored, namely the trajectory record Tr i and the hidden state H i , where H i can be split into four parts, namely and In the initialization stage, Tr i has one observation, and the values of H i are all set to 0, that is

[0087] S3. Use the prediction model to perform a single-step prediction on the tracking trajectory of the tracker and enter the next round of matching, including the following steps:

[0088] S3.1. For the i-th tracker, take out the vehicle coordinate points at the latest two moments from the trajectory record and calculate the relative displacement of their latitudes and longitudes:

[0089]

[0090] where, and respectively represent the vehicle coordinate points of the trajectory record of the i-th tracker at time t - 1;

[0091] S3.2. Concatenate the relative displacements of the latitudes and longitudes, which is the overall displacement of the vehicle represented by the i-th tracker at time t - 1 Take it as the input of the long short-term memory network:

[0092]

[0093]

[0094] where, ‖ represents the concatenation operation; L-LSTM is the first LSTM in the present invention, called the local LSTM, which is used to encode the time embedding of a single vehicle; and respectively represent the hidden state and cell state of the i-th tracker corresponding to the L-LSTM network at time t; W l and b l respectively represent the weights and biases of the L-LSTM, which are shared by all vehicles at the same time;

[0095] S3.3. Then, take as the input of the graph attention network to obtain the attention coefficients of the implicit spatial interaction between vehicles:

[0096]

[0097] where is the attention coefficient of node j to node i at time t; · T represents matrix transpose; represents the set of neighbor nodes of node i on the graph; W ∈ R F′×F is the weight matrix shared by all nodes, used to perform a linear transformation on the features of each node, where F represents the dimension of, and F′ is the dimension of the output; a ∈ R 2F′ is the shared attention mechanism used to calculate the attention coefficients; a undergoes matrix operations and is then normalized through the activation function LeackyReLU and the Softmax method;

[0098] S3.4. Use the attention coefficients for calculation to obtain the spatial embedding of the vehicle represented by the i-th tracker

[0099]

[0100] where represents concatenating K spatial attention embeddings; σ(·) represents the activation function, which is a non-linear method; is the normalized attention coefficient calculated by the k-th attention mechanism, W k is the weight matrix of the corresponding input linear transformation;

[0101] S3.5. Take as the input of the second long short-term network:

[0102]

[0103] where G-LSTM is the second LSTM in the present invention, called the global LSTM, which is used to encode the time embedding of each vehicle after being implicitly spatially influenced by other vehicles; and respectively represent the hidden state and cell state of the i-th tracker corresponding to the G-LSTM network at time t; W g and bg represent the weights and biases of the G-LSTM respectively, and are shared by all vehicles at the same moment;

[0104] S3.6. Add the local encoding and the global encoding of each vehicle to obtain the output embedding of the vehicle represented by each tracker At this time is equivalent to the residual; then use this input embedding as the input of the fully connected layer to obtain the final result:

[0105]

[0106] where FC(·) represents the fully connected layer; W f and b f are the weights and biases of the fully connected layer respectively;

[0107] S3.7. After obtaining the relative displacement prediction of each tracker at time t add it to the vehicle coordinate point of the tracker at time t-1 to obtain the prediction value of the tracker at time t

[0108] S4. When training the prediction model of the present invention, use the improved deep Hungarian model to further calculate the prediction result of the prediction model as the tracking accuracy, and realize the end-to-end training with the tracking accuracy as the goal. S4 includes the following steps:

[0109] S4.1. In the training stage, first calculate the D matrix using a method similar to that shown in S2.1:

[0110]

[0111] where α is a scaling factor, and the value of the calculation result needs to be scaled to the range of [0, 0.1]. This value needs to be selected according to the specific application scenario;

[0112] S4.2. Use D as the input of the deep Hungarian model to obtain the soft assignment matrix The size of this matrix is the same as the input, and the output value is between [0, 1];

[0113] S4.3. Given the correct matches, the true positive matrix B of D can be obtained TP , as well as the true positive one-dimensional matrices P r and P c :

[0114]

[0115]

[0116]

[0117] Among them, represents the value of B TP at the i-th row and j-th column. If the predicted value of the i-th tracker and the observed value of the j-th vehicle are a true positive match, then this value is 1, otherwise it is 0;

[0118] S4.3. According to the method of the deep Hungarian model, for add a row (or column) of data δ as a threshold in the row (or column) dimension respectively, and then calculate Softmax in the row (or column) dimension to obtain C r (or C c ); The present invention further processes it to obtain improved false positives and false negatives

[0119]

[0120] S4.4: According to the method of the deep Hungarian model, the representing ID switching errors can be calculated, and then these three errors are calculated as tracking metrics MOTA and MOTP as loss functions:

[0121]

[0122]

[0123] Among them, M represents the number of vehicle observations; γ is a control coefficient used to adjust 's importance; ⊙ represents the Hadamard product; ‖·‖1 represents the sum of all values in the matrix, and ‖·‖0 represents the number of all non-zero values in the matrix; λ is a control coefficient used to adjust the importance of dMOTP.

Claims

1. A vehicle trajectory tracking method based on a spatio-temporal interaction neural network, characterized in that, It includes the following steps: S1. Preprocess the sensor data, convert the relative coordinates to geographical coordinates, and cluster to obtain the average geographical coordinate points; S2. Match the vehicle observations with the tracker prediction values and update the tracker trajectory record; S3. Use the prediction model to perform a single-step prediction on the tracking trajectory of the tracker and enter the next round of matching; S4. Use the improved deep Hungarian network to perform end-to-end training on the prediction model with the tracking accuracy as the goal; The specific content of step S1 includes: S1.

1. Denote the relative coordinates obtained by the sensor as where represents the i-th coordinate obtained by the j-th sensor; convert it to an approximate value of the geographical coordinates wherein and represent the longitude and latitude of the geographical coordinates obtained by converting the vehicle; lng0 and lat0 represent the longitude and latitude of the geographical coordinates of the sensor, serving as the coordinate origin; R represents the radius of the earth, in meters; S1.

2. Cluster the geographical coordinates after conversion of the sensors on all the same road sections with a distance of half a meter using the distance-based clustering algorithm, and then take the average value within the cluster: Among them, represents the average in-cluster geographical coordinates of the clustering cluster c and serves as the representative of this clustering cluster; is the number of in-cluster geographical coordinates of the clustering cluster c; and respectively represent the longitude and latitude of the i-th point within the clustering cluster c; S1.

3. Convert the transformed geographical coordinates of all sensors on the road network into clusters using a distance-based clustering algorithm at a distance of half a meter, and take the in-cluster average using Equation (3); finally, obtain the processed data coordinate point set P t , representing all vehicle observations obtained at time t; The specific content of step S2 includes the following steps: S2.

1. Calculate for P t and the predicted values of the tracker Calculate the coordinate distances pairwise: where d i,j represents the predicted coordinate point of the i-th tracker at time t and the coordinate point of the j-th vehicle at the same time the distance between; D is a matrix of size representing the number of trackers at time t representing the number of vehicle observations at time t; D i,j represents the value of the i-th row and j-th column in the matrix, when d i,j ≤ 1 meter, take d i,j , otherwise take Inf to represent infinity; S2.

2. Calculate D using the Hungarian algorithm to obtain the minimum-cost matching between the tracker and the vehicle observations; S2.

3. For the vehicle observations with matching objects, add them to the matching tracker as the trajectory record of the tracker; for the vehicle observations without matching objects, generate a new tracker for each observation and add it to the tracker; for the tracker without a matching object, add its prediction at the current moment to the trajectory record. If the consecutive unmatched time of the tracker exceeds the threshold ε, delete the tracker; S2.

4. For each tracker, two contents are stored, namely the trajectory record Tr i and the hidden state H i , where H i can be split into four parts, namely and In the initialization stage, Tr i has one observation value, while the values of H i are all set to 0, that is The specific content of step S3 includes the following steps: S3.

1. For the i-th tracker, take out the vehicle coordinate points at the latest two moments from the trajectory record and calculate the relative displacement of their longitude and latitude; Among them, and respectively represent the vehicle coordinate points of the i-th tracker trajectory record at the (t - 1) moment; S3.

2. Concatenate the relative displacements of longitude and latitude, which is the overall displacement of the vehicle represented by the i-th tracker at time t-1 Use it as the input to the long short-term memory network: Among them, ‖ represents the concatenation operation; L-LSTM is the first LSTM, called the local LSTM, which is used to encode the temporal embedding of a single vehicle; and respectively represent the hidden state and cell state of the L-LSTM network corresponding to the i-th tracker at time t; W l and b l respectively represent the weights and biases of the L-LSTM, which are shared by all vehicles at the same moment; S3.

3. Then, take as the input of the graph attention network to obtain the attention coefficients of the implicit spatial interaction between vehicles: Among them, is the attention coefficient of node j to node i at time t; · T represents matrix transpose; represents the set of neighbor nodes of node i on the graph; W ∈ R F′×F is the weight matrix shared by all nodes, used to perform a linear transformation on the features of each node, where F represents the dimension of, and F′ is the output dimension; a ∈ R 2F′ is the shared attention mechanism used to calculate the attention coefficient; a undergoes matrix operations and is then normalized through the activation function LeackyReLU and the Softmax method; S3.

4. Calculate using the attention coefficient to obtain the spatial embedding of the vehicle represented by the i-th tracker Among them, denotes the concatenation of K spatial attention embeddings; σ(·) denotes the activation function, which is a non-linear method; is the normalized attention coefficient calculated by the k-th attention mechanism, and W k is the weight matrix of the corresponding input linear transformation; S3.

5. Take as the input of the second long short-term memory network: Among them, G-LSTM is the second LSTM, called the global LSTM, which is used to encode the temporal embedding of each vehicle after being affected by the implicit space of other vehicles; and respectively represent the hidden state and cell state of the i-th tracker corresponding to the G-LSTM network at time t; W g and b g respectively represent the weights and biases of G-LSTM, which are shared by all vehicles at the same moment; S3.

6. Add the local encoding and the global encoding of each vehicle to obtain the output embedding of the vehicle represented by each tracker At this time equivalent to the residual; then use this input embedding as the input of the fully connected layer to obtain the final result: Among them, FC(·) represents the fully connected layer; W f and b f are the weight and bias of the fully connected layer, respectively; S3.

7. After obtaining the relative displacement prediction of each tracker at time t add it to the vehicle coordinate point of the tracker at time t - 1 to obtain the predicted value of the tracker at time t The specific content of step S4 includes: S4.

1. In the training stage, first calculate the D matrix: where α is the scaling coefficient, and the numerical value of the calculation result needs to be scaled to the range of [0, 0.1]. This value needs to be selected according to the specific application scenario; S4.

2. Use D as the input of the deep Hungarian model to obtain the soft assignment matrix The size of this matrix is the same as the input, and the values are between [0, 1]; S4.

3. In the case of a known correct match, the true positive matrix B of D can be obtained TP , as well as the true positive one-dimensional matrix P of the true positive matrix in the row and column dimensions r and P c : Among them, represents B TP the value at the i-th row and j-th column. If the predicted value of the i-th tracker and the observed value of the j-th vehicle are a true positive match, then this value is 1; otherwise, it is 0. S4.

3. According to the method of the deep Hungarian model, for add a column of data δ as a threshold in the column dimension respectively, and then calculate Softmax in the row dimension to obtain C r ; for add a row of data δ as a threshold in the row dimension, and then calculate Softmax in the column dimension to obtain C c ; further process it to obtain the improved false positives and false negatives S4.

4. According to the method of the deep Hungarian model, the Then, these three errors are calculated as the tracking metrics dMOTA and dMOTP as the loss function.

2. The vehicle trajectory tracking method based on a spatio-temporal interaction neural network according to claim 1, wherein: The error of the approximation value of the geographical coordinates described in step S1.1 does not exceed 0.1 meter.