Track representation learning method based on fuzzy coding
Through the pyramid structure encoder-decoder model (BLUE model) based on fuzzy encoding, the spatio-temporal resolution, generalization ability and robustness of the existing trajectory representation learning method when processing GPS trajectories is solved, and more efficient trajectory representation and multi-level semantic learning are achieved.
Patent Information
- Application Number
- CN202510295458.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
AI Technical Summary
The existing trajectory representation learning methods have limited spatial and temporal resolution, poor generalization ability and poor robustness when dealing with GPS trajectories, and it is difficult to effectively capture the fine-grained spatial and temporal details of the trajectory and generalize in different cities.
A trajectory representation learning method based on fuzzy encoding is proposed. Using the encoder-decoder model (BLUE model) of pyramid structure, the GPS trajectory is compressed into patch trajectory through fuzzy encoding, and hierarchical semantic learning and trajectory reconstruction are performed through attention mechanism and cross attention.
This method is better than the existing method in terms of effectiveness and efficiency, and can better capture the multi-level semantics of trajectories and improve the performance of tasks such as trajectory classification and similarity search.
Smart Images

Figure CN120218157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trajectory processing, and in particular to a method for learning trajectory representation based on fuzzy coding. Background Art
[0002] With the popularization of GPS devices, the movement trajectories of individuals and vehicles can be easily collected. Analyzing these trajectories is crucial for many urban tasks, such as traffic management, commercial services, and regional planning, which contribute to the construction of smart cities. As a fundamental task of trajectory analysis, trajectory representation learning generates a vector representation for each trajectory while preserving the spatio-temporal travel semantics. By converting variable-length trajectories into fixed-length vectors, trajectory representation learning supports many downstream tasks, such as trajectory classification, travel time estimation, and trajectory similarity calculation. For example, trajectory similarity calculation initially required quadratic time for calculating the trajectory length, but it becomes linear-time Euclidean distance calculation using the vectors generated by trajectory representation learning.
[0003] Existing methods use grid trajectories or road trajectories as inputs to handle the irregularity and noise of GPS trajectories. Although grid and road trajectories are beneficial to TRL by capturing high-level travel semantics (i.e., regions or roads), they also have three key limitations. 1. Limited resolution: By grouping multiple GPS points into a grid cell or a road segment, they capture high-level local information but lose the fine-grained spatio-temporal details inherent in individual GPS points, which are crucial for maintaining basic spatio-temporal properties. In particular, a grid cell represents a geographical area, so an exact timestamp cannot be assigned to it. Similarly, a road segment usually only retains the average time of its GPS points. The limited spatio-temporal resolution hinders tasks that require fine-grained information (e.g., travel time estimation). 2. Poor generalization ability: The grid cells and road networks are different in different cities, so models using grid or road trajectories cannot generalize across different cities. However, good generalization ability is beneficial for reducing the training cost of the model and handling the situation of scarce data. 3. Poor robustness: Grid-based TRL methods need to adjust the size of grid cells to obtain good accuracy. The quality of road trajectories depends on the map matching algorithm and the accuracy of the road map, so the accuracy of road-based methods is also affected. Summary of the Invention
[0004] In order to solve the above problems, the present invention proposes a method for learning trajectory representation based on fuzzy coding.
[0005] The technical solution of the present invention is: A method for learning trajectory representation based on fuzzy coding includes the following steps:
[0006] S1. Collect the original GPS trajectory, and use the encoder of the BLUE model to encode the original GPS trajectory to obtain the final input embedding of the GPS points;
[0007] S2. Generate a compressed trajectory representation based on the final input embedding of the GPS points;
[0008] S3. Perform trajectory reconstruction using the decoder of the BLUE model based on the compressed trajectory representation;
[0009] S4. Determine the spatio-temporal loss of the trajectory reconstruction and train the BLUE model;
[0010] S5. Generate a trajectory representation using the trained BLUE model.
[0011] Furthermore, S1 includes the following sub-steps:
[0012] S11. Determine the spatial context vector of the GPS points according to the forward distance, forward azimuth, backward distance, and backward azimuth of the GPS points in the original GPS trajectory, and convert it into a spatial encoding using a linear layer;
[0013] S12. Convert the temporal context vector of the GPS points in the original GPS trajectory into a temporal encoding;
[0014] S13. Obtain the final input embedding of the GPS points according to the spatial encoding and temporal encoding of the GPS points in the GPS trajectory.
[0015] Furthermore, in S11, the spatial encoding of the GPS points is expressed as:
[0016] ;
[0017] In the formula, represents the spatial context vector of the GPS points, represents the linear layer.
[0018] In S12, the temporal encoding of the GPS points is expressed as:
[0019] ;
[0020] In the formula, represents the first parameter for learning temporal context information, represents the second parameter for learning temporal context information, represents the temporal context vector of the GPS points, represents the concatenation operation;
[0021] In S13, the final input embedding of the GPS points The expression is:
[0022] .
[0023] Furthermore, S2 includes the following sub-steps:
[0024] S21. Add positional encoding to the final input embedding of the GPS points and generate the first-level basic spatio-temporal information;
[0025] S22. Convert the first-level basic spatio-temporal information into the first-level patch trajectory;
[0026] S23. Calculate the attention scores of the first-level patch trajectory embeddings and calculate the normalized attention scores based on the attention scores of the patch trajectory embeddings;
[0027] S24. Calculate the second-level patch trajectory embeddings based on the normalized attention scores;
[0028] S25. Repeat S23 - S24 to obtain the third-level patch trajectory embeddings, and generate the compressed trajectory representation based on the first-level patch trajectory embeddings, the second-level patch trajectory embeddings, and the third-level patch trajectory embeddings.
[0029] Furthermore, in S21, the first-level basic spatio-temporal information The expression is:
[0030] ;
[0031] ;
[0032] In the formula, represents the Transformer layer, represents the representation vector of the current-level global mutual relationship learned after the TransformerEncoder, represents the first GPS point, represents the second GPS point, represents the last GPS point of the trajectory, represents the length of the trajectory;
[0033] In S22, the first-level patch trajectory The expression is:
[0034] ;
[0035] ;
[0036] ;
[0037] In the formula, represents aligning different patch lengths to the same patch length using padding embedding, represents by using to embed GPS points into patches, represents the trajectory embedding containing the first-level basic spatio-temporal information and global information, represents the patch length of the first-level basic spatio-temporal information, represents the representation vector of the first GPS point of the trajectory, represents the representation vector of the second GPS point, represents the representation vector of the last GPS point of the trajectory, represents the first scalar, represents the second scalar, represents the scalar, represents the length of the patch trajectory of the next level;
[0038] In S23, the attention score of the patch trajectory embedding is expressed as:
[0039] ;
[0040] In the formula, represents the patch trajectory embedding, represents the multi-layer perceptron network;
[0041] In S23, the normalized attention score is expressed as:
[0042] ;
[0043] In the formula, represents the exponent, represents the th attention score corresponding to the representation vector of the
[0044] In S24, the patch trajectory embedding of the second level is expressed as:
[0045] ;
[0046] ;
[0047] In the formula, The representation vector indicating the first position of the second-level patch trajectory The representation vector indicating the second position of the second-level patch trajectory The representation vector indicating the last position of the second-level patch trajectory
[0048] Furthermore, S3 includes the following sub-steps:
[0049] S31. Embed and restore the patch trajectory of the third level to the patch trajectory embedding of the second level;
[0050] S32. Calculate the global information based on the restored patch trajectory embedding of the second level;
[0051] S33. Based on the global information, restore the patch trajectory embedding of the second level to the patch trajectory embedding of the first level to complete the trajectory reconstruction.
[0052] Furthermore, in S31, the restored patch trajectory embedding of the second level has the following expression:
[0053] ;
[0054] In the formula, represents the cross-attention network, represents the query token, represents the key token, represents the value token, represents the trajectory representation vector after being encoded by the TransformerEncoder at the second level in the encoder network, represents the input trajectory representation vector of the third level in the decoder network;
[0055] In S32, the global information has the following expression:
[0056] ;
[0057] In the formula, represents the self-attention network;
[0058] In S33, the restored patch trajectory embedding of the first level has the following expression:
[0059] ;
[0060] In the formula, represents the network architecture of the Transformer, It represents the trajectory characterization vector after the second layer in the encoder network passes through the Transformer Encoder.
[0061] Furthermore, S4 includes the following sub-steps:
[0062] S41. Calculate the spatial reconstruction loss and the temporal reconstruction loss;
[0063] S42. Calculate the spatio-temporal loss according to the spatial reconstruction loss and the temporal reconstruction loss;
[0064] S42. Perform an averaging operation on the spatio-temporal losses of all GPS trajectories to obtain the reconstruction loss, and complete the training of the BLUE model.
[0065] Furthermore, in S41, the spatial reconstruction loss has the following expression:
[0066] ;
[0067] In the formula, represents the characterization vector of the reconstructed GPS point in the trajectory , represents a multi-layer perceptron network;
[0068] In S41, the temporal reconstruction loss has the following expression:
[0069] ;
[0070] In S42, the spatio-temporal loss of the trajectory has the following calculation formula:
[0071] ;
[0072] In the formula, represents the trajectory, represents the length of the trajectory, represents the spatial attribute vector of the position in the trajectory, represents the temporal attribute vector of the position in the trajectory.
[0073] The beneficial effects of the present invention are as follows: The present invention proposes a trajectory representation learning method based on fuzzy coding. The constructed BLUE model is an encoder-decoder model with a pyramid structure. The encoder compresses the GPS trajectory into a patch trajectory from low level to high level, and the decoder reconstructs the GPS trajectory from high level to low level. The BLUE model based on the pyramid structure of the present invention can learn the hierarchical semantics of multi-level patch trajectories, and is superior to the existing trajectory representation learning methods in terms of effectiveness and efficiency, promoting tasks such as trajectory classification and similarity search. Description of the Drawings
[0074] Figure 1 It is a flowchart of a trajectory representation learning method based on fuzzy coding;
[0075] Figure 2 It is a schematic diagram of the overall architecture of the BLUE model;
[0076] Figure 3 It is a diagram of the fuzzy coding process;
[0077] Figure 4 It is a diagram of the execution process of the encoder;
[0078] Figure 5 It is a comparison diagram of training time and inference time;
[0079] Figure 6 It is a diagram of the dimension experiment results;
[0080] Figure 7 It is a diagram of the layer number experiment results. Detailed implementation manners
[0081] The following further describes the embodiments of the present invention with reference to the accompanying drawings.
[0082] As Figure 1 shown, the present invention provides a trajectory representation learning method based on fuzzy coding, including the following steps:
[0083] S1. Collect the original GPS trajectory, and use the encoder of the BLUE model to encode the original GPS trajectory to obtain the final input embedding of the GPS points;
[0084] S2. Generate a compressed trajectory representation according to the final input embedding of the GPS points;
[0085] S3. Perform trajectory reconstruction using the decoder of the BLUE model according to the compressed trajectory representation;
[0086] S4. Determine the spatio-temporal loss of the trajectory reconstruction, and train the BLUE model;
[0087] S5. Generate a trajectory representation using the trained BLUE model.
[0088] Based on the patch trajectory, the present invention designs a model architecture of an encoder and a decoder with a pyramid structure. The encoder gradually compresses and transforms the low-level GPS trajectory into a high-level patch trajectory with multiple model levels. Each level contains two components, namely a Transformer for capturing the global information of the trajectory at the current level, and a patch pooling for capturing the local semantics of the patch blocks and generating a higher-level patch trajectory. On the contrary, the decoder gradually reconstructs the low-level GPS trajectory from the high-level patch trajectory. Each level of the decoder contains an upsampling to recover the current-level trajectory from a higher level, and a Transformer to refine the recovered trajectory to improve the accuracy. To handle the variable patch lengths in patch pooling and different trajectory lengths during upsampling, the present invention designs an attention-based patch pooling method and uses cross-attention for upsampling. To train the BLUE model, the mean square error loss is used to reconstruct the original GPS trajectory, which is more effective than the cross-entropy loss (involving Softmax operations for a large number of classes) used in grid-based and road-based methods.
[0089] Define GPS trajectory: GPS trajectory is a time-series sequence of GPS points collected by a GPS device at fixed or non-fixed time intervals. In it, each GPS point is where , and represent longitude, latitude, and timestamp respectively.
[0090] Define trajectory representation learning: Given a set of GPS trajectories, trajectory representation learning aims to learn a -dimensional vector representation for each trajectory in the set. The learned representation is expected to achieve high accuracy for downstream tasks such as travel time estimation, trajectory classification, and search for the most similar trajectories.
[0091] Figure 2 Shows the overall structure of BLUE, which is a pyramid-structured encoder-decoder model composed of four key modules: Spatiotemporal embedding: Converts spatiotemporal GPS points into hidden embeddings; Patch encoder: Learns hierarchical spatiotemporal embeddings by gradually reducing GPS accuracy to create patches from high resolution to low resolution; Block decoder: Uses cross-attention for length-insensitive recovery to recover high-resolution trajectories from compressed low-resolution blocks; Spatiotemporal reconstruction: Trains the model by reconstructing spatiotemporal GPS trajectories with mean square error loss.
[0092] The present invention proposes a trajectory-specific patch method based on different GPS decimal precisions, which determines hierarchical patches by gradually discarding the least significant decimals of GPS coordinates. Block trajectories are generated to learn high-level travel semantics, and GPS trajectories are retained to learn detailed low-level spatio-temporal semantics. As Figure 3 shown, the details of the fuzzy encoding are illustrated. Specifically, three precision levels are considered: precision@5 (level-1) with a 1m resolution, representing the original GPS trajectory; precision@3 (level-2) with a 100m resolution, and precision@2 (level-3) with a 1km resolution, representing patch trajectories; Precision@4 (10m resolution) is omitted because it has no significant difference from precision@5. To transition from precision@5 to precision@3, a rounding method is used to group GPS points with similar spatial semantics at precision@5 into the same patch at precision@3. This process converts the GPS trajectory into a patch trajectory. For example, , and are grouped into the same patch at precision@3. This grouping process is repeated at precision@2 to generate new block trajectories that capture higher-level travel semantics. The input remains precision@5 because the lower precisions are not used directly but are inferred through these mappings.
[0093] The present invention utilizes the semantic information carried by GPS points to blur the GPS resolution and group low-level points into small blocks. Therefore, it has the following advantages: 1. Adaptive patch length, the algorithm dynamically adjusts the precision of the trajectory block length based on GPS semantics and density, enabling it to adapt to trajectories of different lengths. For example, a trajectory with many clustered GPS points will have a shorter patch trajectory length, while a trajectory with fewer and more widely spaced points may have a patch length of 1. 2. Semantic consistency, GPS points within a region have consistent low-resolution semantics, representing meaningful local attributes, reducing spatio-temporal inconsistencies in the original GPS trajectory and avoiding truncated motion semantics. 3. Hierarchical representation, trajectories at different levels reveal unique information, reflecting hierarchical semantics; for example, the first level, i.e., precision@5, retains basic spatio-temporal details; the second level, i.e., precision@3, captures local motion behaviors such as turns; the third level, i.e., precision@2, provides insights into long-range overall patterns. In addition, the length of the block trajectories at the higher levels is also greatly reduced, thereby improving efficiency.
[0094] Compared with patches in computer vision and time series prediction, fuzzy coding solves two problems in trajectories: 1. Variable length. In computer vision, using a fixed patch size of 64×64, a batch of images with the same resolution of 512×512 can be converted into a patch sequence of length 64; similarly, in time series prediction, using a patch size of 8, a batch sequence of fixed length 128 can be transformed into a patch sequence of length 16. However, trajectory batches vary greatly in length, from a few points to hundreds of points, so it is impractical to define a fixed patch size for them. 2. Local semantics. In computer vision, a patch captures local spatial semantics, while in time series prediction, a fixed patch length (e.g., points based on a 5-minute sampling rate with a length of 12) represents the time semantics of one hour. However, for trajectories, it is difficult for a fixed block length to capture meaningful motion behaviors, usually splitting dense clusters or continuous points and disturbing the motion semantics.
[0095] In an embodiment of the present invention, S1 includes the following sub-steps:
[0096] S11. Determine the spatial context vector of a GPS point according to the forward distance, forward azimuth angle, backward distance, and backward azimuth angle of the GPS point in the original GPS trajectory, and convert it into a spatial encoding using a linear layer;
[0097] S12. Convert the time context vector of the GPS point in the original GPS trajectory into a time encoding;
[0098] S13. Obtain the final input embedding of the GPS point according to the spatial encoding and time encoding of the GPS point in the GPS trajectory.
[0099] The spatio-temporal embedding layer aims to convert the coordinates and timestamps of the GPS trajectory, i.e., the first layer, into spatio-temporal hidden embeddings. The present invention designs two encoding modules to achieve this, namely spatial encoding and time encoding.
[0100] Spatial encoding: The original GPS point has only limited spatial information, i.e., coordinates. To enrich the spatial information, additional spatial relationships are introduced within the trajectory. Specifically, for a GPS point , calculate its forward distance and forward azimuth angle to the next GPS point , and calculate its backward distance and backward azimuth angle to the previous GPS point . If is the first point of the trajectory, the backward distance and backward azimuth angle are zero. If is the last point of the trajectory, the forward distance and forward azimuth angle are zero. Therefore, combining the longitude and latitude information, Spatial context vector , and then use a linear layer to convert it into a spatial encoding.
[0101] Temporal encoding: The original GPS time information of the GPS point is a timestamp that includes year, month, day, hour, minute, and second. However, the gaps between these values are large. For example, the period of a year is usually 365, and the period of a second is usually 60, resulting in inconsistent ranges of the input data. To solve this problem, the present invention uses day-of-year, day-of-month, day-of-week, hour-of-day, minute-of-hour, and second-of-minute to convert each value into the same range [0, 1], and further subtracts 0.5 to keep the range of each value [-0.5, 0.5]. Thus, a temporal context vector of a GPS point is obtained , and then it is converted into a temporal information encoding.
[0102] In an embodiment of the present invention, in S11, the spatial encoding of the GPS point has the following expression:
[0103] ;
[0104] In the formula, represents the spatial context vector of the GPS point, represents the linear layer.
[0105] In S12, the temporal encoding of the GPS point has the following expression:
[0106] ;
[0107] In the formula, represents the first parameter for learning temporal context information, represents the second parameter for learning temporal context information, represents the temporal context vector of the GPS point, represents the concatenation operation;
[0108] In S13, the final input embedding of the GPS point has the following expression:
[0109] .
[0110] In an embodiment of the present invention, S2 includes the following sub-steps:
[0111] S21. Add positional encoding to the final input embedding of the GPS points and generate the first-level basic spatio-temporal information;
[0112] S22. Convert the first-level basic spatio-temporal information into the patch trajectories of the first level;
[0113] S23. Calculate the attention scores of the patch trajectory embeddings of the first level and calculate the normalized attention scores according to the attention scores of the patch trajectory embeddings;
[0114] S24. Calculate the patch trajectory embeddings of the second level according to the normalized attention scores;
[0115] S25. Repeat S23 - S24 to obtain the patch trajectory embeddings of the third level and generate the compressed trajectory representation according to the patch trajectory embeddings of the first level, the patch trajectory embeddings of the second level, and the patch trajectory embeddings of the third level.
[0116] The backbone network of the present invention is an encoder-decoder model with a pyramid structure. The encoder compresses the trajectory through fuzzy encoding to learn the trajectory representation from the low level to the high level. The decoder aims to recover the low-level trajectory from the compressed high-level trajectory.
[0117] Patch Encoder: The patch encoder alternates between Transformer and fuzzy encoding. Figure 4 Shows the execution process of the encoder. Transformer is first used to capture the global sequence information of the current level, and then fuzzy encoding is performed through patch and pooling to compress the trajectory embedding, generate the patch embedding with local semantics, and use it as the input of the next layer. Specifically, first add positional encoding to the GPS trajectory embedding and then use Transformer to learn the basic spatio-temporal information of the first layer, and then fuzzy it by converting it into patch trajectories.
[0118] To generate the next-level trajectory embedding through the patch embedding the present invention proposes an attention-based pooling method to eliminate the influence of the padding embedding caused by the operation and provide more effective pooling results. Softmax is used to normalize the attention scores within a patch.
[0119] Repeat the above process to obtain the globally relevant trajectory embedding passing through Transformer in the second level and obtain the new patch trajectory input in the third level Then, execute the Transformer again to learn the final compressed trajectory representation , to capture high-level travel semantics and global information. At the upper layer, since the sequence length is reduced, positional encoding is still applied before each use of the Transformer.
[0120] In the embodiment of the present invention, in S21, the first-level basic spatio-temporal information has the following expression:
[0121] ;
[0122] ;
[0123] In the formula, represents the Transformer layer, represents the representation vector of the global mutual relationship learned at the current level after the Transformer Encoder, represents the first GPS point, represents the second GPS point, represents the last GPS point of the trajectory, represents the length of the trajectory;
[0124] In S22, the patch trajectory of the first level has the following expression:
[0125] ;
[0126] ;
[0127] ;
[0128] In the formula, represents using padding embedding to align different patch lengths to the same patch length, represents passing through using to generate patches from GPS point embeddings, represents the trajectory embedding containing the first-level basic spatio-temporal information and global information, represents the patch length of the first-level basic spatio-temporal information, represents the representation vector of the first GPS point of the trajectory, represents the representation vector of the second GPS point, represents the representation vector of the last GPS point of the trajectory, represents the first scalar, represents the second scalar, represents the scalar, Indicates the length of the patch trajectory at the next level;
[0129] In S23, the attention score of the patch trajectory embedding has the following expression:
[0130] ;
[0131] In the formula, represents the patch trajectory embedding, represents a multi-layer perceptron network; MLP consists of a linear layer, a normalization layer, a ReLU activation function, and another linear layer, which are executed sequentially. If is calculated by padding the embedding , set to float(-inf). Softmax normalization is used for the attention scores within a patch; if is float(-inf), then .
[0132] In S23, the normalized attention score has the following expression:
[0133] ;
[0134] In the formula, represents the exponent, represents the th attention score corresponding to the th representation vector in the th patch,
[0135] In S24, the patch trajectory embedding at the second level has the following expression:
[0136] ;
[0137] ;
[0138] In the formula, represents the representation vector at the first position of the patch trajectory at the second level, represents the representation vector at the second position of the patch trajectory at the second level, represents the representation vector at the last position of the patch trajectory at the second level, ; represents the representation vector at the th position of the patch trajectory at the second level.
[0139] In an embodiment of the present invention, S3 includes the following sub-steps:
[0140] S31. Embed and restore the patch trajectory of the third level to the patch trajectory embedding of the second level;
[0141] S32. Calculate the global information according to the restored patch trajectory embedding of the second level;
[0142] S33. According to the global information, restore the patch trajectory embedding of the second level to the patch trajectory embedding of the first level to complete the trajectory reconstruction.
[0143] Patch decoder: The patch decoder is symmetric in structure with the patch encoder and alternates between upsampling operations and Transformer. The upsampling operation aims to restore the next lower-level trajectory from the current high-level trajectory. The Transformer is used to capture the global correlation in the restored sequence and enhance the reconstruction ability. First, a Transformer is applied to derive a new trajectory embedding from the final trajectory representation to ensure symmetry in the model architecture and a deeper patch decoder structure.
[0144] Restoring the next lower-level trajectory from the current high-level using the upsampling operation is the opposite of the patch pooling process in the patch encoder. The compressed trajectory embedding of the patch decoder is used as the input, and the trajectory embedding of the patch encoder at the resolution corresponding to the expected restored trajectory is used to assist the restoration process. As the length of the restored sequence increases, the self-attention mechanism is applied to enable the low-level trajectory representation to capture the global information over the extended length. Again, using the of the patch encoder as a reference, it is added to the restored trajectory embedding as the input of the Transformer as a shortcut connection for more accurate trajectory reconstruction at the second level.
[0145] The above decoder process can use the of the patch decoder and the of the patch encoder to obtain the trajectory embedding at the first level . Here, the positional encoding is discarded in the Transformer because and carry positional information in the patch encoder.
[0146] In an embodiment of the present invention, in S31, the expression of the restored patch trajectory embedding of the second level is:
[0147] ;
[0148] Wherein, represents the cross-attention network, represents the query token, represents the key token, represents the value token, represents the trajectory representation vector after being encoded by the TransformerEncoder in the second layer of the encoder network, represents the input trajectory representation vector of the third layer in the decoder network; is the query, which comes from the output of the second layer Transformer of the patch encoder; Key and value are respectively , which comes from the output of the third layer Transformer of the patch decoder. Taking as the query is because the output shape of the cross-attention is consistent with the shape of the query, and the length of is exactly the length that wants to be recovered from only contributes to the calculation of the cross-attention matrix. The actually recovered trajectory embedding is calculated by weighted summation with , ensuring that is only used for auxiliary reconstruction.
[0149] In S32, the global information has the following expression:
[0150] ;
[0151] Wherein, represents the self-attention network;
[0152] In S33, the expression of the recovered patch trajectory embedding of the first layer is:
[0153] ;
[0154] Wherein, represents the network architecture of the Transformer, represents the trajectory representation vector after the second layer of the encoder network passes through the TransformerEncoder.
[0155] In the embodiment of the present invention, S4 includes the following sub-steps:
[0156] S41. Calculate the spatial reconstruction loss and the temporal reconstruction loss;
[0157] S42. Calculate the spatio-temporal loss based on the spatial reconstruction loss and the temporal reconstruction loss;
[0158] S42. Average the spatio-temporal losses of all GPS trajectories to obtain the reconstruction loss, and complete the training of the BLUE model.
[0159] The present invention incorporates the temporal loss to improve the reconstruction accuracy of low-level trajectories. In addition, compared with the mask model loss used in grid-based and road-based methods, the mean square loss is more efficient because the mask model loss relies on cross-entropy classification. Considering that there are usually tens of thousands of road segments and grids, the Softmax operation in cross-entropy becomes very inefficient.
[0160] In the embodiment of the present invention, in S41, the spatial reconstruction loss has the following expression:
[0161] ;
[0162] In the formula, represents the representation vector of the reconstructed GPS point in the trajectory , represents the multi-layer perceptron network;
[0163] In S41, the temporal reconstruction loss has the following expression:
[0164] ;
[0165] In S42, the spatio-temporal loss of the trajectory has the following calculation formula:
[0166] ;
[0167] In the formula, represents the trajectory, represents the length of the trajectory, represents the spatial attribute vector of the position in the trajectory, represents the temporal attribute vector of the position in the trajectory.
[0168] Next, the BLUE model of the present invention will be tested and described.
[0169] The time complexity of the BLUE model mainly comes from the attention mechanism of the Transformer, where is the sequence length. Let the length of the first-level trajectory be , and the number of layers of the first-level Transformer be ; the length of the second-level trajectory is , the number of layers of the second - level Transformer ; the length of the third - level trajectory is , the number of layers of the third - level Transformer is . The training time complexity of the BLUE model, including the patch encoder and the patch decoder , where and are cross - attentions, and are self - attentions. The inference time complexity of the BLUE model only includes the patch encoder, which is , where . Due to the reduction of the trajectory length, the actual operation becomes very efficient, that is, approximately . Table 1 shows the average sequence lengths of different types of trajectories. The average sequence length of the patch trajectories of BLUE is significantly reduced, and the average length of the third level is in single digits on both datasets. Experiments show that BLUE has an advantage in terms of efficiency compared with the state - of - the - art trajectory representation learning methods.
[0170] Table 1
[0171]
[0172] The present invention has been experimented on two real - world trajectory datasets, namely Porto and Chengdu, which have been widely used in previous trajectory representation learning research. The ratios of training data, validation data, and test data for both datasets are set to [0.6, 0.2, 0.2]. The statistical data of the datasets are shown in Table 2. The GPS points include latitude, longitude, and timestamp, and are sampled every 15 seconds and 30 seconds on average in Porto and Chengdu respectively. Porto records 3 different vehicle travel modes, including starting from the central administrative area, directly asking the taxi driver to stop at a specific stand, and a trip on any street. Chengdu includes whether the taxi is carrying passengers. After pre - processing, the present invention obtains 1133593 and 1114194 trajectories in Porto and Chengdu respectively.
[0173] Table 2
[0174]
[0175] The present invention compares BLUE with eight state-of-the-art trajectory representation learning methods, namely Traj2vec, TrajCL, Trembr, PIM, JCLRNT, START, MMTEC, and JGRM. Among them, TrajCL performs best in grid-based methods, START performs best in road-based methods, and JGRM performs best in multi-modal methods.
[0176] The hidden embedding and trajectory representation dimension of the BLUE model are set to 128. The Transformer configurations of the patch encoder and decoder include: the first layer has 2 layers of Transformer, the second layer has 4 layers of Transformer, and the third layer has 2 layers of Transformer. The attention mechanism uses 4 heads, and the dropout is 0.1. The present invention uses the Adam optimizer to pre-train BLUE, with a batch size of 256 and a learning rate of 1e-4. All methods are trained for 30 epochs, and the model of the last epoch is used for evaluation. According to the dataset statistics, the maximum patch length M is 10 (from the first layer to the second layer) and 40 (from the second layer to the third layer) for the Porto dataset, and 10 and 45 for the Chengdu dataset.
[0177] The present invention selects three downstream tasks widely used in trajectory representation learning: Travel Time Estimation, Trajectory Classification, and Most Similar Trajectory Search. The present invention does not take trajectory recovery or trajectory prediction as downstream tasks because grid-based and road-based methods regard these as classification problems, while GPS-based methods regard them as regression problems. Therefore, their evaluation metrics are different, making direct comparison unreasonable. Following existing trajectory representation learning work, the present invention selects the same evaluation metrics for these tasks. For travel time estimation, the present invention reports the Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and Root Mean Square Error (RMSE). For multi-classification of the Porto dataset, the present invention reports Micro-F1 and Macro-F1. For binary classification of the Chengdu dataset, the present invention reports the F1 value, accuracy, and precision. For the most similar trajectory search, the present invention reports the hit rate of the top-k results, namely HR@1, HR@5, and Mean Rank (MR).
[0178] Tables 3 and 4 compare BLUE with eight other trajectory representation learning methods. The improvements of BLUE over the best-performing baselines are generally substantial. For travel time estimation, the performance gains range from a minimum of 50.66% to a maximum of 88.31%. On the Porto dataset, the proposed travel time estimation achieves single-digit MAE (in seconds) and a very low MAPE of only 1%. Different from road- and grid-based methods that group multiple GPS points into road segments or grid cells, resulting in the loss of important spatio-temporal details, the underlying layer of BLUE preserves the complete GPS information, thus having superior performance. Compared with the Porto dataset, the performance on the Chengdu dataset is slightly lower due to the highly uneven sampling rate, which ranges from 3 seconds to 60 seconds. In contrast, the sampling rate of the Porto dataset is typically around 15 seconds, providing more consistency.
[0179] For the most similar trajectory search, BLUE achieves improvements from 3.36% to over 37.99%. Although the spatio-temporal irregularity of GPS trajectories usually affects the performance of similarity search, the BLUE model alleviates this problem by converting GPS trajectories into patch trajectories through fuzzy coding and its pyramid structure. This process reduces the spatio-temporal irregularity, while the high-level trajectory representation effectively compresses the trajectory information and captures hierarchical regional patterns using low-level fine-grained details, thus significantly improving the similarity search performance.
[0180] For trajectory classification, compared with other tasks, the improvement of the BLUE model is less obvious because classification is relatively easier and mainly relies on the overall trajectory semantics (such as OD and travel patterns), which rely less on fine-grained details. In addition, the Chengdu dataset is a binary classification and maintains high performance in all methods, which can be reflected from the accuracy of the strong baseline methods.
[0181] Table 3
[0182]
[0183] Table 4
[0184]
[0185] To verify the design of BLUE, the present invention conducted experiments on its six variants. w / o P@2 deletes precision@2, and the final trajectory representation is the [CLS] of precision@3. w / o P@3 deletes precision@3. The patch trajectory of precision@2 comes from precision@5. w / o P@5 deletes precision@5. Here, we only delete the Transformer of precision@5 to keep the input as a GPS trajectory. w / Min replaces the attention pooling with Min pooling. w / Max replaces the attention pooling with Max pooling. w / Mean replaces the attention pooling with Mean pooling.
[0186] Table 5 shows the results of the ablation study. The results indicate that all designs effectively improve the model performance, as deleting any component leads to a performance drop. Generally speaking, the impact on the trajectory similarity search (MSTS) is the most significant, as it directly depends on the quality of the pre-trained model. In contrast, tasks such as travel time estimation (TTE) and trajectory classification (TC) experience less performance degradation due to fine-tuning. Specifically, the low-level representation, i.e., P@5, has the most substantial impact on the task performance, as it captures the basic spatio-temporal correlations within the trajectory through the Transformer. On the other hand, the higher levels have relatively less impact on the overall performance. This is because the higher levels mainly rely on the lower-level modules, where the semantic information of the trajectory is severely compressed. Additionally, different pooling strategies also significantly affect the model performance. The Min pooling and Max pooling methods focus on extracting the most prominent features in the spatio-temporal information of the trajectory. The Mean pooling preserves more spatio-temporal integrity of the trajectory but results in feature smoothing, which may lead to the loss of spatio-temporal details.
[0187] Table 5
[0188]
[0189] The present invention conducted experiments to evaluate the portability of the BLUE model and SOTA methods. The results are shown in Table 6, indicating a significant performance decline for all baseline models. This is mainly because the road network scales and grid scales in different cities are different, resulting in the need to re-initialize the grid id and road id for road-based and grid-based methods. For the BLUE model, the performance loss for all tasks is the smallest. This is because the BLUE model only relies on unified GPS and time information as input without any external dependencies, making the model highly versatile. It is worth noting that the BLUE model can be effectively migrated for tasks such as MSTS without fine-tuning. However, when migrating from Porto to Chengdu, since the city scale of Porto is smaller than that of Chengdu, the data distribution fails to cover the entire range of trajectories in Chengdu, resulting in a significant MR loss. On the contrary, the transfer performance from Chengdu to Porto in MSTS exceeds the result without transfer, indicating that models trained on larger cities may produce better results when transferred to smaller cities.
[0190] Table 6
[0191]
[0192] The present invention evaluated the efficiency of the model using three metrics: model size, training time, and inference time. The results are shown in Table 7 and Figure 5 as follows, demonstrating the high efficiency of the model of the present invention. Since the model of the present invention has no external dependencies, the model size is not affected by the city scale. In contrast, the model parameters of road-based and grid-based models increase significantly with the increase in city scale. Despite incorporating 20 attention layers, the BLUE model achieves fast training time and effective inference, which is largely attributed to its pyramid structure, significantly reducing the sequence length at higher levels. It is worth noting that BLUE inference is very fast because it only requires an encoder. In addition, compared with the mask recovery loss of the Softmax operation adopted by START and JGRM, the mean squared error loss used by the BLUE model in training is more computationally efficient. TrajCL benefits from its lightweight network, namely a two-layer network; MMTEC uses NeuralCDE and a complex network structure design, requiring a longer training time.
[0193] Table 7
[0194]
[0195] The present invention conducted experiments on different hidden embedding dimensions of the final trajectory representation, ranging from [16, 32, 64, 128, 256]. Figure 6Shows the results of two tasks: travel time estimation (fine-tuning) and most similar trajectory search (without fine-tuning) on two datasets. Generally, both tasks improve with the increase of the embedding dimension, because it allows for a more comprehensive encoding of travel semantics in the hidden embeddings and the final trajectory representations. Specifically, when the Chengdu dataset increases from 128 dimensions to 256 dimensions, the travel time estimation accuracy improves, while for the Porto dataset, the accuracy decreases. This shows that the dimensionality trend varies across different datasets, but overall, the 128-dimensional embedding achieves the best balance between performance and efficiency.
[0196] The present invention changes the number of layers of each level of the Transformer: for the first and third levels, the number of layers varies between [1, 2, 3], and for the second level, it varies between [2, 4, 6]. The present invention has conducted experiments on the Chengdu dataset, and the results are as Figure 7 shown. The overall trends of the first and third levels are similar. When the number of layers is set to 1 or 3, the performance will decline. This is because too few layers are unable to effectively learn trajectory information. For the first level, due to limited GPS trajectory information, too many layers will lead to overfitting. For the third level, it also does not require too many layers because the length of the patch sequence is greatly reduced. However, the second level, as the intermediate layer, constructed by the first level and supporting the third level, requires more layers. Therefore, it needs a deeper module for effective representation. It is worth noting that when the number of layers is set to 4, the performance of the third level stabilizes.
[0197] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A trajectory representation learning method based on fuzzy coding, characterized in that: The following steps are involved: S1, collect the original GPS trajectory, encode the original GPS trajectory using the encoder of the BLUE model, and obtain the final input embedding of the GPS point; S2, generate a compressed trajectory representation based on the final input embedding of the GPS points; S3, based on the compressed trajectory representation, the trajectory is reconstructed using the decoder of the BLUE model; S4, determine the spatiotemporal loss of trajectory reconstruction and train the BLUE model; S5. Generate trajectory representation using the trained BLUE model.
2. The trajectory representation learning method based on fuzzy coding according to claim 1 is characterized in that: The S1 comprises the following sub-steps: S11, determining the spatial context vector of the GPS point according to the forward distance, forward azimuth, backward distance and backward azimuth of the GPS point in the original GPS trajectory, and converting it into a spatial encoding using a linear layer; S12, converting the time context vector of the GPS point in the original GPS trajectory into a time code; S13. Obtain the final input embedding of the GPS point according to the spatial coding and temporal coding of the GPS point in the GPS trajectory.
3. The trajectory representation learning method based on fuzzy coding according to claim 2 is characterized in that: In S11, the spatial encoding of the GPS point The expression is: ; In the formula, represents the spatial context vector of the GPS point, represents a linear layer, In S12, the time code of the GPS point The expression is: ; In the formula, represents the first parameter used to learn the temporal context information, represents the second parameter used to learn temporal context information, represents the temporal context vector of the GPS point, Represents a splicing operation; In S13, the final input of the GPS point is embedded The expression is: 。 4. The trajectory representation learning method based on fuzzy coding according to claim 1 is characterized in that: The S2 comprises the following sub-steps: S21, adding location encoding to the final input embedding of the GPS point and generating the first level of basic spatiotemporal information; S22, converting the first-level basic spatiotemporal information into a first-level patch track; S23, calculating the attention score of the first-level patch trajectory embedding, and calculating the normalized attention score according to the attention score of the patch trajectory embedding; S24, calculate the second-level patch trajectory embedding according to the normalized attention score; S25. Repeat S23-S24 to obtain the patch trajectory embedding of the third level, and generate a compressed trajectory representation according to the patch trajectory embedding of the first level, the patch trajectory embedding of the second level, and the patch trajectory embedding of the third level.
5. The trajectory representation learning method based on fuzzy coding according to claim 4 is characterized in that: In S21, the first level basic spatiotemporal information The expression is: ; ; In the formula, represents the Transformer layer, Represents the representation vector of the global mutual relationship of the current level learned after TransformerEncoder, Indicates the first GPS point, Indicates the second GPS point, Indicates the last GPS point of the track. represents the length of the trajectory; In S22, the first-level patch track The expression is: ; ; ; In the formula, It means that different patch lengths are aligned to the same patch length using padding embedding. Indicates that by using Embed GPS points into patches. represents the trajectory embedding containing the first-level basic spatiotemporal information and global information, The patch length representing the first-level basic spatiotemporal information, The characterization vector representing the first GPS point of the trajectory, The characterization vector representing the second GPS point, The characterization vector representing the last GPS point of the trajectory, represents the first scalar, represents the second scalar, Indicates Scalar, Indicates the length of the next level patch track; In S23, the attention score of the patch trajectory embedding The expression is: ; In the formula, represents the patch trajectory embedding, represents a multilayer perceptron network; In S23, the normalized attention score The expression is: ; In the formula, represents the index, Indicates The first patch The attention score corresponding to the representation vector, Indicates the maximum patch length; In S24, the second-level patch track is embedded The expression is: ; In the formula, Represents the representation vector of the first position of the second-level patch track, Represents the representation vector of the second position of the second-level patch track, A representation vector representing the final position of the second-level patch track.
6. The trajectory representation learning method based on fuzzy coding according to claim 1 is characterized in that: The S3 comprises the following sub-steps: S31, restoring the patch trajectory embedding of the third level to the patch trajectory embedding of the second level; S32, calculating global information according to the restored patch trajectory embedding of the second level; S33: According to the global information, the patch trajectory embedding of the second level is restored to the patch trajectory embedding of the first level to complete the trajectory reconstruction.
7. The trajectory representation learning method based on fuzzy coding according to claim 6 is characterized in that: In S31, the second-level restored patch trajectory is embedded The expression is: ; In the formula, represents the cross attention network, Indicates the query symbol. Indicates a keyword token, Indicates the value symbol, Represents the trajectory representation vector of the second level in the encoder network after being encoded by TransformerEncoder, represents the input trajectory representation vector of the third layer in the decoder network; In S32, global information The expression is: ; In the formula, represents the self-attention network; In S33, the first level of restored patch trajectory is embedded The expression is: ; In the formula, Represents the Transformer network architecture, Represents the trajectory representation vector of the second level in the encoder network after passing through TransformerEncoder.
8. The trajectory representation learning method based on fuzzy coding according to claim 1 is characterized in that: The S4 comprises the following sub-steps: S41, calculating the spatial reconstruction loss and the temporal reconstruction loss; S42, calculating the space-time loss according to the space reconstruction loss and the time reconstruction loss; S42, averaging the spatiotemporal losses of all GPS trajectories to obtain the reconstruction loss, and completing the training of the BLUE model.
9. The trajectory representation learning method based on fuzzy coding according to claim 8 is characterized in that: In S41, the spatial reconstruction loss The expression is: ; In the formula, Representation trajectory The characterization vector of the reconstructed GPS point, represents a multi-layer perceptron network; In S41, the time reconstruction loss The expression is: ; In S42, the space-time loss of the trajectory The calculation formula is: ; In the formula, Represents the trajectory, represents the length of the trajectory, is the spatial attribute vector representing the position in the trajectory, A vector of time attributes representing positions in the trajectory.
Citation Information
Cited By
Time sequence prediction method based on multi-modal contrast learning technology
CN121457754A