A vehicle trajectory spatio-temporal reconstruction method and system based on an attention mechanism
Through the vehicle trajectory spatiotemporal reconstruction method based on the attention mechanism, the problem of sparse trajectory reconstruction is solved by utilizing trajectory embedding and multi-head self-attention mechanism combined with long short-term memory network, achieving efficient and accurate trajectory reconstruction and improving the performance of the intelligent transportation system.
Patent Information
- Application Number
- CN202310578357.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Sparse automatic vehicle recognition trajectories lead to the loss of details in the trajectories, affecting the application effect of intelligent transportation systems. Existing technologies find it difficult to effectively reconstruct sparse vehicle trajectories.
A vehicle trajectory spatiotemporal reconstruction method based on attention mechanism is adopted to reconstruct the vehicle's spatiotemporal trajectory through trajectory embedding, historical trajectory aggregation and multi-head self-attention mechanism, combined with long short-term memory network.
It achieves accurate and efficient reconstruction of sparse trajectories, improves the granularity and application value of trajectories, and enhances the accuracy and efficiency of intelligent transportation systems.
Smart Images

Figure CN116578661B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of spatiotemporal reconstruction of automatic vehicle recognition trajectories, and relates to a vehicle trajectory spatiotemporal reconstruction method and system based on an attention mechanism. Background Art
[0002] Automatic Vehicle Identification (AVI) technology is widely used in urban traffic monitoring. Its accurate and comprehensive trajectory data has greatly promoted the development and application of intelligent transportation systems. Compared with traditional trajectory data, the greatest advantage of AVI data is that it can cover almost all vehicles traveling in urban traffic. At the same time, it also has faster recognition speed and more accurate recognition capabilities. However, due to economic and environmental constraints, the layout of AVI monitoring points in cities is very sparse and unevenly distributed, resulting in sparse spatiotemporal positions within AVI trajectories. This introduces great uncertainty within the trajectories, which greatly undermines the usefulness of AVI trajectories. Therefore, reconstructing sparse AVI trajectories into fine-grained trajectories has become a key issue in AVI trajectory mining, with important theoretical and practical significance.
[0003] Compared to traditional GPS tracks, automatic vehicle identification (AVI) data offers many advantages. First, AVI data is highly accurate. This high accuracy stems from two factors. First, because AVI data is collected at specific locations, it eliminates uncertainty about the location of the track. Second, it can accurately identify vehicles and record their travel times. AVI systems, in particular, based on radio frequency identification (RFID), can achieve near-100% accuracy for vehicle identification. Second, AVI data is massive. This volume is reflected in three key aspects. First, AVI systems can identify all types of vehicles in a city, including private cars, taxis, ride-hailing services, and trucks. This is the system's greatest advantage, as it covers virtually all types of vehicles. Second, AVI technology can be applied in a variety of scenarios, such as toll booths and urban roads. Third, it is rich in information. AVI data captures more information, such as vehicle ID, vehicle type, and vehicle usage. Therefore, the data volume of AVI systems far exceeds that of GPS, covering scenarios and vehicles.
[0004] However, automatic vehicle identification data also has some drawbacks. The most significant one is its low sampling rate. This is because automatic vehicle identification equipment cannot record vehicle locations in real time; it can only record vehicle information as it passes through collection points. Due to environmental and economic constraints, collection points in urban road networks are often sparsely distributed, with the distance between two adjacent collection points ranging from a few hundred meters to several kilometers. This directly leads to very sparse vehicle trajectories based on automatic vehicle identification data. Existing research has shown that sparse vehicle trajectories can lead to the loss of details in the trajectories, severely impacting subsequent research and applications. Therefore, if automatic vehicle identification trajectories are to fully leverage their advantages and role in the construction of intelligent transportation systems, they must be reconstructed. Summary of the Invention
[0005] In view of this, the object of the present invention is to provide a vehicle trajectory spatiotemporal reconstruction method and system based on an attention mechanism, so as to accurately reconstruct the vehicle's spatiotemporal trajectory.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for spatiotemporal reconstruction of vehicle trajectories based on an attention mechanism is proposed. The method comprises the following steps: S1, embedding the spatial trajectory of the current trajectory into a trajectory; S2, embedding multiple historical trajectories into a trajectory; S3, using a time series prediction model based on the attention mechanism to predict the time series corresponding to the current trajectory and combine them into a spatiotemporal trajectory; S4, using real data to train the overall model and reconstruct the trajectory in a test set. Finally, the prediction error is evaluated and analyzed based on the prediction results and actual data.
[0008] Furthermore, in step S1, the spatial trajectory of the current trajectory is embedded in the trajectory, which specifically includes: the processing of the current trajectory is divided into two parts, namely, spatial reconstruction of the current trajectory and trajectory embedding of the reconstructed spatial trajectory; for a line traveling between adjacent point pairs (C A ,C B ) in the trajectory [(as A ,at A ),(as B ,as B )], according to the reconstruction result of its spatial trajectory, its spatial trajectory is represented as an m-dimensional vector S = [s1, s2, ..., s m ], where s i represents the i-th spatial point in the spatial trajectory;
[0009] For the current trajectory, use a time state vector S tThe total travel time characteristic is achieved by converting the total travel time (at A -at B ) is expanded into an m-dimensional vector (with the same dimension as the space vector S), S t The calculation method is as follows:
[0010] S t =(at B -at A )×S
[0011] In order to express the influence of vehicle departure time on the spatiotemporal relationship, a period vector is used to represent the vehicle driving cycle, and the time of a day is divided into C time slices, and then C S To indicate the departure time of the trip at A If the time slice is within a day, the period vector of the trip can be expressed as S C , and can be obtained by the following formula:
[0012]
[0013] According to the current trajectory embedding method mentioned above, the embedded current trajectory C can be expressed as:
[0014] C=Concat(S,S t ,S c ).
[0015] Furthermore, in step S2, multiple historical trajectories are embedded and expressed. For a given pair of adjacent collection points (C A ,C B ), whose historical trajectory set is represented by T H ={Tr1,Tr2,…,Tr N}, considering one of the trajectories Tr i , the density of its internal spatiotemporal positions is much lower than the density of the path space points. The trajectory aggregation method is used to aggregate all historical trajectories into a complete spatiotemporal trajectory, including:
[0016] S21. Historical trajectory aggregation:
[0017] The method of the present invention is similar to that of the previous one, except that the spatial position in the path is divided into spatial slices instead of time slices. Then, for a certain spatial slice, all the time points in the time slice are found, and the average time of these trajectory points is used as the time of the aggregated trajectory in the spatial slice. Therefore, for the historical trajectory set TH , the aggregated trajectory It can be expressed as follows:
[0018]
[0019] where Tr i Represents the historical trajectory set T H The i-th trajectory in the trajectory set is N, where N represents the number of trajectories in the trajectory set and ⊕ represents the trajectory aggregation operation. Specifically, for a point at a fixed position, find the time of all points passing through the position and use the average of the time spent passing through the position as the time of the point. After aggregating the time for each position point, sort the positions in position order to obtain the aggregated trajectory.
[0020] S22. Historical trajectory aggregation of the same period:
[0021] The historical trajectory of the same period is defined as follows: Given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N}, for a trajectory to be reconstructed [(as A ,at A ),(as B ,at B )], find all the trajectory departure time at A The historical trajectories with the same period are the collection of historical trajectories with the same period, which is expressed as T C ={Tr1,Tr2,…,Tr M}, where m represents the number of trajectories in the set, Tr i Represents a track in the set, and aggregates the historical tracks of the same period. The operation formula is as follows:
[0022]
[0023] S23. Historical trajectory aggregation of the same state:
[0024] The historical trajectory of the same state is defined as follows: Given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N};For a trajectory to be reconstructed [(as A ,at A ),(as B ,at B )], first according to the travel time (atB -at A ) calculate the traffic status of the trip, and then for T H All trajectories in , calculate the traffic status of each trajectory according to its travel time, and represent the set of trajectories with the same state as the trajectory to be reconstructed as T S ={Tr1,Tr2,…,Tr O}, where O represents the number of trajectories in the set, Tr i Represents a trajectory in the set, and aggregates the historical trajectories of the same state. The operation is shown in the following formula, where ⊕ represents trajectory aggregation:
[0025]
[0026] S24. Historical trajectory fusion:
[0027] For the three aggregated historical trajectories, in order to enable the model to capture the spatiotemporal relationships within and between trajectories, a splicing operation is used to fuse the trajectories and input H as the historical trajectory of the model. The specific operation is as follows:
[0028]
[0029] From the three historical trajectories used above, we can see that the historical trajectory set contains a historical trajectory set with the same period and a historical trajectory set with the same state. The reason why they are extracted separately is to strengthen the weights of these two trajectories in the neural network, because the trajectory to be reconstructed may have a higher similarity with these two trajectories.
[0030] Furthermore, in step S3, in order to reconstruct the time series of the current spatial trajectory, the spatiotemporal relationship within the trajectory is learned from the historical trajectory, that is, the input of the model is the historical spatiotemporal trajectory and the current spatial trajectory, and the output is the time series corresponding to the current spatial trajectory; for this time series generation task, an encoder-decoder architecture is adopted, in which the encoder is responsible for processing the historical spatiotemporal trajectory, and the model uses a multi-head self-attention mechanism to encode the historical spatiotemporal trajectory and extract the spatiotemporal relationship in the historical trajectory; in the decoder, the spatial trajectory is input, and the multi-head self-attention mechanism is still used to extract the spatial features therein, and cross attention is used to combine the encoder and decoder to learn the spatiotemporal relationship within the trajectory from the output of the encoder; at the end of the model, a long short-term memory network is used to capture the temporal relationship between trajectories, and finally the required time series is output.
[0031] Further, the model is divided into an encoder and a decoder, the encoder is responsible for encoding the historical trajectory, and the decoder is responsible for processing the current trajectory and generating a time series jointly with the output of the encoder, specifically including:
[0032] Encoder: the input of the encoder is the historical trajectory H, in order to enable the encoder to better extract the features therein, a linear layer is used to embed the historical trajectory for expression, and map it to a higher-dimensional hidden space:
[0033] H h =Linear(H)
[0034] Where H h is the output of the linear layer; in the space-time trajectory, the space-time relationship is not only related to the state of the trajectory, but also related to the time period in which it is located, and the space-time relationship is highly related to the total time it passes through the path, therefore, a multi-head attention mechanism is used to capture the space-time relationship, first map the input sequence H h to Q h , K h and V h three subspaces:
[0035] Q H =W Q H h
[0036] K H =W K H h
[0037] V H =W V H h
[0038] Where W Q , W K , W V are the weight matrices of Q, K, V, i.e. mapping H h to Q, K, V three hidden spaces; for this, the operation of multi-head self-attention is defined as:
[0039]
[0040] MultiHead(Q,K,V)=[Head1,Head2,...,Head h ]W O
[0041] Head i =Attention(QW i Q ,KW i K,VW i V )
[0042] Where h is the number of attention heads, W O is the final output weight matrix; after the multi-head self-attention, the model then uses the residual connection and layer normalization layer. The specific formula is as follows:
[0043] LayerNorm(X+MultiHeadAttention(X))
[0044] The residual connection refers to X + MultiHeadAttention (X), where X refers to the input of the multi-head self-attention. LayerNorm refers to layer normalization, which converts the mean and variance of the input of neurons in each layer to a uniform value. The multi-head self-attention and residual connection & layer normalization modules in the encoder are repeated L times. The output of the previous layer is used as the input of the next layer until the output of the last layer of the encoder. The attention formula is then used to map the output of the last layer to the two latent spaces K and V for query by the decoder.
[0045] Decoder: The decoder takes the current trajectory C as input, which contains the current trajectory's spatial information, state information, period information, and travel time information. To make it easier for the decoder to extract this information, a linear layer is used to embed the current trajectory C and map it into a higher-dimensional latent space. The specific operation is as follows:
[0046] C h =Linear(C)
[0047] C h is the output of the linear layer. After embedding the trajectory, a two-layer self-attention mechanism is used in the decoder to capture the spatiotemporal relationship between and within trajectories. In the first layer’s self-attention module, residual connection, and layer normalization, the processing of spatial trajectories is consistent with the self-attention mechanism in the encoder.
[0048] In the second-layer multi-head self-attention module, a cross-attention mechanism is used to ensure that the encoder output has different values depending on the decoder input. In this layer, the encoder's final output is mapped into two latent spaces, K and V, and the output of the previous decoder layer is mapped into the latent space Q. A multi-head self-attention operation is then performed on the mapped Q, K, and V. To address model training issues and accelerate model convergence, the output of the cross-attention mechanism is processed using a residual connection and layer normalization. Furthermore, the encoder uses a multi-layer structure, with the output of the previous decoder layer serving as the input to the next decoder layer, and so on until the output of the final layer.
[0049] Finally, in order to enable the model to capture the temporal and spatial relationships between trajectories, a long short-term memory network is used to process the decoder output. Before that, in order to enable the LSTM to better capture the temporal relationship between trajectories, the decoder output is merged with the original spatial trajectory, that is, concatenated. The final LSTM output is the required time series:
[0050] T=LSTM(Concat( dec ,C))
[0051] Where T represents the final generated time series, O dec represents the output of the decoder, and C represents the spatial trajectory after embedding;
[0052] After generating the time series, we only need to concatenate it with the unembedded spatial trajectory bit by bit to obtain the reconstructed spatiotemporal trajectory.
[0053] The beneficial effects of the present invention are:
[0054] The vehicle trajectory spatiotemporal reconstruction method and system based on the attention mechanism provided by the present invention can accurately and efficiently reconstruct the vehicle's spatiotemporal trajectory. Compared with existing technical solutions, the technical solution of the present invention has higher accuracy and efficiency, and this solution has broad application prospects.
[0055] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0057] Figure 1 It is a model diagram of the present invention;
[0058] Figure 2 This is a time series prediction module diagram of the present invention;
[0059] Figure 3 It is an experimental effect diagram of the present invention. DETAILED DESCRIPTION
[0060] The technical solution of the present invention is described in detail below with reference to the accompanying drawings.
[0061] Figure 1 This is a model diagram of the present invention. As shown in the figure, the present invention provides a vehicle trajectory spatiotemporal reconstruction method based on the attention mechanism, which mainly includes the following steps: S1. Embedding the spatial trajectory of the current trajectory; S2. Embedding multiple historical trajectories; S3. Using a time series prediction model based on the attention mechanism to predict the time series corresponding to the current trajectory and combine them into a spatiotemporal trajectory; S4. Using real data to train the overall model and reconstruct the trajectory in the test set. Finally, the prediction error is evaluated and analyzed based on the prediction results and actual data.
[0062] In this embodiment, step S1 embeds the current trajectory into a trajectory. In the model of the present invention, the processing of the current trajectory is very important, because the original current trajectory only contains the spatiotemporal position information of the starting point and the ending point. It is very difficult to generate the information of countless points in the middle of the trajectory based on the information of these two points. Therefore, the spatial trajectory is reconstructed first, and then the spatiotemporal trajectory is reconstructed based on the spatial trajectory. Therefore, the processing of the current trajectory is divided into two parts, namely, spatial reconstruction of the current trajectory and trajectory embedding of the reconstructed spatial trajectory; for a vehicle traveling between adjacent points (C A ,C B ) in the trajectory [(as A ,at A ),(as B ,at B )], according to the reconstruction result of its spatial trajectory, its spatial trajectory is represented as an m-dimensional vector S = [s1, s2, ..., s m ], where s i represents the i-th spatial point in the spatial trajectory;
[0063] For the current trajectory, use a time state vector S t The total travel time characteristic is achieved by converting the total travel time (at A -at B ) is expanded into an m-dimensional vector (with the same dimension as the space vector S), St The calculation method is as follows:
[0064] S t = (at B -at A ) x S
[0065] In order to represent the influence of the vehicle departure time on the space-time relationship, a cycle vector is used to represent the vehicle travel cycle, the time of a day is divided into C time slices, and then C S is used to represent the departure time at of the trip A in the time slice position within a day, and the cycle vector of the trip can be represented as S C , and can be obtained by the following formula:
[0066]
[0067] According to the above-mentioned current trajectory embedding method, the embedded current trajectory C can be represented as:
[0068] C = Concat (S, S t , S c ).
[0069] In step S2, a plurality of historical trajectories are embedded and expressed, for a given pair of adjacent collection points (C A , C B ) between a certain path, the historical trajectory set is represented as T H = {Tr1, Tr2, …, Tr N}, considering one of the trajectories Tr i , the density of the internal space-time position is much lower than the density of the path space point, using the trajectory aggregation method to aggregate all historical trajectories into a complete space-time trajectory, specifically including:
[0070] S21, historical trajectory aggregation:
[0071] By separating the space-time positions in the trajectory according to the time slice, and then taking the maximum probability position in each time slice as the position of the aggregated trajectory on the time slice. The method of the present application is similar, the difference is that the present application divides the space position in the path according to the space slice instead of the time slice, and then for a certain space slice, find all the time points on the time slice, use the average time of these trajectory points as the time of the aggregated trajectory on the space slice; therefore, for the historical trajectory set T H , the aggregated trajectory can be represented as follows:
[0072]
[0073] where Tr i Represents the historical trajectory set T H The i-th trajectory in the trajectory set is N, where N represents the number of trajectories in the trajectory set and ⊕ represents the trajectory aggregation operation. Specifically, for a point at a fixed position, find the time of all points passing through the position and use the average of the time spent passing through the position as the time of the point. After aggregating the time for each position point, sort the positions in position order to obtain the aggregated trajectory.
[0074] S22. Historical trajectory aggregation of the same period:
[0075] Considering the influence of the periodicity of the trajectory on the path on the vehicle motion law, the historical trajectories of the same period are extracted from the historical trajectories, and the aggregated trajectories are also used as the input of the model; the historical trajectories of the same period are defined as follows: given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N}, for a trajectory to be reconstructed [(as A ,at A ),(as B ,at B )], find all the trajectory departure time at A The historical trajectories with the same period are the collection of historical trajectories with the same period, which is expressed as T C ={Tr1,Tr2,…,Tr M}, where m represents the number of trajectories in the set, Tr i Represents a track in the set, and aggregates the historical tracks of the same period. The operation formula is as follows:
[0076]
[0077] S23. Historical trajectory aggregation of the same state:
[0078] Considering that the regularity of vehicle movement on the road is affected by the traffic state (fast passing / congested), the traffic state of the vehicle, i.e., congested or fast passing, can be calculated based on the time it takes for the vehicle to move between two adjacent points. When the vehicle's travel time is greater than 1.5 times the average time, the trip is considered congested. The historical trajectory of the same state is defined as follows: Given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N};For a trajectory to be reconstructed [(as A,at A ),(as B ,at B )], first according to the travel time (at B -at A ) calculate the traffic status of the trip, and then for T H All trajectories in , calculate the traffic status of each trajectory according to its travel time, and represent the set of trajectories with the same state as the trajectory to be reconstructed as T S ={Tr1,Tr2,…,Tr O}, where O represents the number of trajectories in the set, Tr i Represents a trajectory in the set, and aggregates the historical trajectories of the same state. The operation is shown in the following formula, where ⊕ represents trajectory aggregation:
[0079]
[0080] S24. Historical trajectory fusion:
[0081] For the three aggregated historical trajectories, in order to enable the model to capture the spatiotemporal relationships within and between trajectories, a splicing operation is used to fuse the trajectories and input H as the historical trajectory of the model. The specific operation is as follows:
[0082]
[0083] From the three historical trajectories used above, we can see that the historical trajectory set contains a historical trajectory set with the same period and a historical trajectory set with the same state. The reason why they are extracted separately is to strengthen the weights of these two trajectories in the neural network, because the trajectory to be reconstructed may have a higher similarity with these two trajectories.
[0084] In step S3, the time sequence generation module in the model is a sequence-to-sequence mode, that is, the input is a sequence and the output is another sequence; in the model, in order to reconstruct the time sequence of the current spatial trajectory, the space-time relationship in the trajectory is learned from the historical trajectory, that is, the input of the model is the historical space-time trajectory and the current spatial trajectory, and the output is the time sequence corresponding to the current spatial trajectory; for the time sequence generation task, an encoder-decoder architecture is used, wherein the encoder is responsible for processing the historical space-time trajectory, the model uses a multi-head self-attention mechanism to encode the historical space-time trajectory, and extracts the space-time relationship in the historical trajectory; in the decoder, the input spatial trajectory is still used to extract the spatial features therein by using the multi-head self-attention mechanism, and the cross attention is used in the decoder to jointly encode the encoder and the decoder, so as to learn the space-time relationship in the trajectory from the output of the encoder; a long short-term memory network is used at the end of the model to capture the time relationship between trajectories, and finally the required time sequence is output. Figure 2 A time sequence prediction module of the application is shown in the figure.
[0085] The model is divided into an encoder and a decoder, the encoder is responsible for encoding the historical trajectory, and the decoder is responsible for processing the current trajectory and generating a time sequence jointly with the output of the encoder, and specifically includes:
[0086] Encoder: the input of the encoder is the historical trajectory H, in order to enable the encoder to better extract the features therein, a linear layer is used to embed the historical trajectory for expression, and the historical trajectory is mapped to a higher-dimensional hidden space:
[0087] H h =Linear(H)
[0088] Where H h is the output of the linear layer; in the space-time trajectory, the space-time relationship is related to the state of the trajectory, and is also related to the time period in which the trajectory is located, and the space-time relationship is highly related to the total time of the trajectory through the path, therefore, a multi-head attention mechanism is used to capture the space-time relationship, first, the input sequence H h is mapped to Q h ,K h and V h three subspaces:
[0089] Q H =W Q H h
[0090] K H =W K H h
[0091] V H =WV H h
[0092] Where W Q 、W K 、W V The weight matrices of Q, K, and V are respectively, that is, H h Mapped to three latent spaces Q, K, and V; for this, the multi-head self-attention operation is defined as:
[0093]
[0094] MultiHead(Q,K,V)=[Head1,Head2,...,Head h ]W O
[0095] Head i =Attention(QW i Q ,KW i K ,VW i V )
[0096] Where h is the number of attention heads, W O is the final output weight matrix; multiple attention heads create multiple subspaces, map the input data to multiple subspaces, and then jointly focus on the representation information of these subspaces to effectively capture the spatiotemporal relationship information in the trajectory;
[0097] After multi-head self-attention, the model then uses residual connections and layer normalization layers to solve the multi-layer network training problem and accelerate convergence. The specific formula is as follows:
[0098] LayerNorm(X+MultiHeadAttention(X))
[0099] The residual connection refers to X+MultiHeadAttention(X), where X refers to the input of the multi-head self-attention. Residual connections are often used in residual networks to allow the network to focus only on the current difference. They are usually used to solve the problem of multi-layer network training. This layer is used here to solve the problem of model training. LayerNorm refers to layer normalization, which converts the mean and variance of the input of neurons in each layer to a uniform value, which can accelerate convergence. The multi-head self-attention and residual connection & layer normalization modules in the encoder are repeated L times. The output of the previous layer is used as the input of the next layer until the output of the last layer of the encoder. The attention formula is then used to map the output of the last layer to the two latent spaces K and V for query by the decoder.
[0100] Decoder: The decoder takes the current trajectory C as input, which contains the current trajectory's spatial information, state information, period information, and travel time information. To make it easier for the decoder to extract this information, a linear layer is used to embed the current trajectory C and map it into a higher-dimensional latent space. The specific operation is as follows:
[0101] C h =Linear(C)
[0102] C h is the output of the linear layer. After embedding the trajectory, a two-layer self-attention mechanism is used in the decoder to capture the spatiotemporal relationship between and within trajectories. In the first layer’s self-attention module, residual connection, and layer normalization, the processing of spatial trajectories is consistent with the self-attention mechanism in the encoder.
[0103] In the second-layer multi-head self-attention module, a cross-attention mechanism is used to ensure that the encoder output has different values depending on the decoder input. In this layer, the encoder's final output is mapped into two latent spaces, K and V, and the output of the previous decoder layer is mapped into the latent space Q. A multi-head self-attention operation is then performed on the mapped Q, K, and V. To address model training issues and accelerate model convergence, the output of the cross-attention mechanism is processed using a residual connection and layer normalization. Furthermore, the encoder uses a multi-layer structure, with the output of the previous decoder layer serving as the input to the next decoder layer, and so on until the output of the final layer.
[0104] Finally, in order to enable the model to capture the temporal and spatial relationships between trajectories, a long short-term memory network is used to process the decoder output. Before that, in order to enable the LSTM to better capture the temporal relationship between trajectories, the decoder output is merged with the original spatial trajectory, that is, concatenated. The final LSTM output is the required time series:
[0105] T=LSTM(Concat( dec ,C))
[0106] Where T represents the final generated time series, O dec represents the output of the decoder, and C represents the spatial trajectory after embedding;
[0107] After generating the time series, we only need to concatenate it with the unembedded spatial trajectory bit by bit to obtain the reconstructed spatiotemporal trajectory.
[0108] The experimental results of this embodiment are as follows:
[0109] Experimental Dataset: We used a real automatic vehicle identification dataset from Chongqing. The dataset contains 330 RFID monitoring points in Chongqing's main urban area. These monitoring points are very unevenly distributed in space, and their adjacent relationships can be learned from their trajectory data. The distance between two adjacent detection points ranges from a few hundred meters to several kilometers. The road structure in between is complex, and the trajectory uncertainty between adjacent points is large. Another dataset contains GPS data from more than 15,000 taxis in Chongqing for one month in October 2021 (2021.10.1-2021.10.31). Chongqing taxi GPS devices record taxi locations every 10-15 seconds. There are approximately 1.5 billion taxi records in this month. After extracting the taxi trips, there are approximately 12 million passenger trips, which are used to train the model.
[0110] Evaluation indicators:
[0111] Spatiotemporal Linear Combine Distance (STLC): The idea behind the spatiotemporal linear combine distance is to split a trajectory into a time series and a space series. When calculating the distance with another trajectory, the distance between the time series and the distance between the space series are considered separately. The time distance and the space distance are then combined according to certain weights to form the overall trajectory distance.
[0112] Spatiotemporal Longest Common Sub-Sequence (STLCSS). Similar to the longest common subsequence (LCSS), but in the spatiotemporal longest common subsequence, the time factor is additionally considered. Originally, only a spatial error range was set. After considering the spatial factor, the temporal error range of the point must also be considered. Both the temporal error and the spatial error must match for the point to be considered a match.
[0113] Experimental results:
[0114] Figure 3 This is the experimental effect diagram of the present invention. Due to the extremely sparse characteristics of AVI trajectories, the spatiotemporal reconstruction of trajectories is divided into two steps in the model of the present invention. First, the spatial features of the trajectory are reconstructed, and then the temporal attributes of the trajectory are reconstructed based on the spatial features of the trajectory. Therefore, in the comparative experiment, the method is also divided into two parts, namely the spatial reconstruction method of the trajectory and the temporal prediction method of the trajectory. The specific method is as follows:
[0115] The spatial reconstruction methods mainly include the following three: RICK, MPR and HMM.
[0116] Time prediction methods mainly include the following: History and Linear belong to traditional statistical learning-based methods, while RNN and LSTM belong to deep learning methods.
[0117] Linear model: A linear model is a simple statistical learning method that uses linear functions to fit the relationship between time and space in a trajectory. That is, for each spatial point, a linear relationship between time and space is assumed. This linear relationship is learned from historical data and used to predict the time of each spatial point.
[0118] History: The history of the most frequent visits is a simple prediction method based on historical data. That is, for any fixed spatial point, the time point with the highest frequency of visits to the spatial point is used as the time of passing the spatial point.
[0119] Recurrent Neural Network (RNN): A recurrent neural network uses a neural network to model the spatiotemporal relationships within the same trajectory. That is, it inputs a spatial sequence of a trajectory and outputs a time sequence of the trajectory. Using a recurrent neural network to model the trajectory can effectively extract the spatiotemporal relationships between different positions within the trajectory.
[0120] Long Short-Term Memory (LSTM) Network: LSTM is a variant of RNN. However, compared with RNN, LSTM Network has the advantage of effectively preserving long-term dependencies in trajectory sequences, and performs better than RNN when the sequence is long.
[0121] Table 1
[0122]
[0123]
[0124] The experimental results are shown in Table 1. Further analysis of the experimental results reveals that, when using a fixed time prediction method and analyzing the impact of different spatial reconstruction methods on spatiotemporal reconstruction, the RICK method performs better than MPR and HMM, while MPR also slightly outperforms the HMM method. Compared to these traditional methods, the present invention achieves superior results, with a significant improvement. When using a fixed spatial reconstruction method and using different time prediction methods for trajectory reconstruction, the History method performs slightly better than the Linear method. Furthermore, the RNN and LSTM methods, which are based on deep learning and take into account spatiotemporal regularities within the trajectory, outperform the History and Linear methods. LSTM also outperforms the RNN method in capturing long-term dependencies within the trajectory, resulting in slightly better experimental results. The proposed time prediction method achieves the best experimental results because it not only considers spatiotemporal dependencies within the trajectory but also learns spatiotemporal relationships from historical trajectories.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A vehicle trajectory spatiotemporal reconstruction method based on an attention mechanism, characterized by: The method comprises the following steps: S1. Embed the spatial trajectory of the current trajectory into a trajectory expression; S2, embedding expressions of multiple historical trajectories; S3. Use a time series prediction model based on the attention mechanism to predict the time series corresponding to the current trajectory and combine them into a spatiotemporal trajectory; S4. Use real data to train the overall model and reconstruct the trajectory in the test set. Finally, evaluate and analyze the prediction error based on the prediction results and actual data. In step S2, the historical trajectory is embedded and expressed. For a given pair of adjacent acquisition points (C A ,C B ), whose historical trajectory set is represented by T H ={Tr1,Tr2,…,Tr N }, considering one of the trajectories Tr i , the density of its internal spatiotemporal positions is much lower than the density of the path space points. The trajectory aggregation method is used to aggregate all historical trajectories into a complete spatiotemporal trajectory, including: S21. Historical trajectory aggregation: The spatial position in the path is divided into spatial slices, and then for a certain spatial slice, all the time points in the time slice are found, and the average time of these trajectory points is used as the time of the aggregated trajectory in the spatial slice; therefore, for the historical trajectory set T H , the aggregated trajectory It can be expressed as follows: where Tr i Represents the historical trajectory set T H The i-th trajectory in , N represents the number of trajectories in the trajectory set, Represents the trajectory aggregation operation. Specifically, for a point at a fixed location, find the time it takes for all points to pass through the location, and use the average of the time it takes to pass through the location as the time of the point. After aggregating the time for each location point, sort the locations in order of location to obtain the aggregated trajectory. S22. Historical trajectory aggregation of the same period: The historical trajectory of the same period is defined as follows: Given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N }, for a trajectory to be reconstructed [(as A ,at A ),(as B ,at B )], find all the trajectory departure time at A The historical trajectories with the same period are the collection of historical trajectories with the same period, which is expressed as T C ={Tr1,Tr2,…,Tr M }, where m represents the number of trajectories in the set, Tr i Represents a track in the set, and aggregates the historical tracks of the same period. The operation formula is as follows: S23. Historical trajectory aggregation of the same state: The historical trajectory of the same state is defined as follows: Given an adjacent point pair (C A ,C B ) and the historical trajectory set T within a path between them H ={Tr1,Tr2,…,Tr N };For a trajectory to be reconstructed [(as A ,at A ),(as B ,at B )], first according to the travel time (at B -at A ) calculate the traffic status of the trip, and then for T H All trajectories in , calculate the traffic status of each trajectory according to its travel time, and represent the set of trajectories with the same state as the trajectory to be reconstructed as T S ={Tr1,Tr2,…,Tr O }, where O represents the number of trajectories in the set, Tr i Represents a trajectory in the set, and aggregates the historical trajectories of the same state. The operation is shown in the following formula, where Aggregation of the represented trajectories: S24. Historical trajectory fusion: For the three aggregated historical trajectories, in order to enable the model to capture the spatiotemporal relationships within and between trajectories, a splicing operation is used to fuse the trajectories and input H as the historical trajectory of the model. The specific operation is as follows: From the three historical trajectories used above, we can see that the historical trajectory set contains a historical trajectory set with the same period and a historical trajectory set with the same state. The reason why they are extracted separately is to strengthen the weights of these two trajectories in the neural network, because the trajectory to be reconstructed has a higher similarity with these two trajectories.
2. The vehicle trajectory spatiotemporal reconstruction method based on the attention mechanism according to claim 1 is characterized by: In step S1, the spatial trajectory of the current trajectory is embedded in the trajectory, specifically including: The processing of the current trajectory is divided into two parts, namely, spatial reconstruction of the current trajectory and trajectory embedding of the reconstructed spatial trajectory; for a line traveling between adjacent points (C A ,C B ) in the trajectory [(as A ,at A ),(as B ,at B )], according to the reconstruction result of its spatial trajectory, its spatial trajectory is represented as an m-dimensional vector S = [s1, s2, ..., s m ], where s i represents the i-th spatial point in the spatial trajectory; For the current trajectory, use a time state vector S t Express the total travel time characteristics, the total travel time (at B -at A ) is expanded into an m-dimensional vector, S t The calculation method is as follows: S t =(at B -at A )×S In order to express the influence of vehicle departure time on the spatiotemporal relationship, a period vector is used to represent the vehicle driving cycle, and the time of a day is divided into C time slices, and then C S To indicate the departure time of the trip at A If the time slice is within a day, the period vector of the trip can be expressed as S C , and can be obtained by the following formula: According to the current trajectory embedding method mentioned above, the embedded current trajectory C can be expressed as: C=Concat(S,S t ,S c )。 3. The vehicle trajectory spatiotemporal reconstruction method based on the attention mechanism according to claim 2 is characterized by: In step S3, in order to reconstruct the time series of the current spatial trajectory, the spatiotemporal relationship within the trajectory is learned from the historical trajectory, that is, the input of the model is the historical spatiotemporal trajectory and the current spatial trajectory, and the output is the time series corresponding to the current spatial trajectory; for this time series generation task, an encoder-decoder architecture is adopted, in which the encoder is responsible for processing the historical spatiotemporal trajectory, and the model uses a multi-head self-attention mechanism to encode the historical spatiotemporal trajectory and extract the spatiotemporal relationship in the historical trajectory; in the decoder, the spatial trajectory is input, and the multi-head self-attention mechanism is still used to extract the spatial features therein, and the cross-attention is used to combine the encoder and decoder to learn the spatiotemporal relationship within the trajectory from the output of the encoder; at the end of the model, a long short-term memory network is used to capture the temporal relationship between trajectories, and finally the required time series is output.
4. The vehicle trajectory spatiotemporal reconstruction method based on the attention mechanism according to claim 3 is characterized by: The model consists of two parts: an encoder and a decoder. The encoder is responsible for encoding historical trajectories, while the decoder is responsible for processing the current trajectory and generating a time series based on the encoder output. Specifically, it includes: Encoder: The input of the encoder is the historical trajectory H, which is embedded using a linear layer and mapped into a higher-dimensional latent space: H h =Linear(H) Among them H h is the output of the linear layer; a multi-head attention mechanism is used to capture its spatiotemporal relationship. First, the input sequence H h Mapping to Q h ,K h and V h Three subspaces: Q H =W Q H h K H =W K H h V H =W V H h Where W Q 、W K 、W V The weight matrices of Q, K, and V are respectively, that is, H h Mapped to three latent spaces Q, K, and V; for this, the multi-head self-attention operation is defined as: MultiHead(Q,K,V)=[Head1,Head2,...,Head h ]W O Head i =Attention(QW i Q ,KW i K ,VW i V ) Where h is the number of attention heads, W O is the final output weight matrix; After the multi-head self-attention, the model then uses residual connections and layer normalization layers. The specific formula is as follows: LayerNorm(X+MultiHeadAttention(X)) The residual connection refers to X + MultiHeadAttention (X), where X refers to the input of the multi-head self-attention. LayerNorm refers to layer normalization, which converts the mean and variance of the input of neurons in each layer to a uniform value. The multi-head self-attention and residual connection & layer normalization modules in the encoder are repeated L times. The output of the previous layer is used as the input of the next layer until the output of the last layer of the encoder. The attention formula is then used to map the output of the last layer to the two latent spaces K and V for query by the decoder. Decoder: The decoder takes the current trajectory C as input, which contains the current trajectory's spatial information, state information, cycle information, and travel time information. A linear layer is used to embed the current trajectory C and map it into a higher-dimensional latent space. The specific operations are as follows: C h =Linear(C) C h is the output of the linear layer. After embedding the trajectory, a two-layer self-attention mechanism is used in the decoder to capture the spatiotemporal relationship between and within trajectories. In the first layer’s self-attention module, residual connection, and layer normalization, the processing of spatial trajectories is consistent with the self-attention mechanism in the encoder. In the second-layer multi-head self-attention module, a cross-attention mechanism is used. In this layer, the final output of the encoder is mapped to two latent spaces, K and V, and the output of the previous decoder layer is mapped to the latent space Q. Then, a multi-head self-attention operation is performed on the mapped Q, K, and V. After that, the output of the cross-attention mechanism is still processed using a residual connection and layer normalization. Finally, a long short-term memory network is used to process the output of the decoder, and the final output of the LSTM is the required time series: T=LSTM(Concat(O dec ,C)) Where T represents the final generated time series, O dec represents the output of the decoder, and C represents the spatial trajectory after embedding; After generating the time series, we only need to concatenate it with the unembedded spatial trajectory bit by bit to obtain the reconstructed spatiotemporal trajectory.
5. A vehicle trajectory spatiotemporal reconstruction system based on an attention mechanism, characterized by: The system adopts the method according to any one of claims 1 to 4.
Citation Information
Cited By
Vehicle control method and system considering early warning effectiveness of operating driver
CN121708770A
A vehicle control method and system considering the effectiveness of the pre-warning of the operating driver
CN121708770B