A trajectory anomaly route detection method and system based on self-supervised trajectory representation learning
By employing a self-supervised trajectory representation learning method, combined with a dual-view synchronous masking strategy using GPS and grid views, and a spatiotemporal fusion mechanism, the problem of insufficient single-view and local anomaly detection in existing trajectory anomaly detection is solved, achieving higher-precision trajectory anomaly detection.
Patent Information
- Application Number
- CN202511164086.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing trajectory anomaly detection methods suffer from low detection accuracy due to the limitations of GPS and grid perspectives, which result in a single, uninterrupted trajectory view and a lack of ability to detect local anomalies.
A self-supervised trajectory representation learning method is adopted. The GPS and grid views are compared and learned through a dual-view synchronous masking strategy. Combined with a spatiotemporal fusion mechanism and Gaussian mixture distribution, local anomalies are captured and reconstruction error is checked.
It improves the sensitivity and detection accuracy of local anomalies, enhances the ability to express trajectory features, reduces the risk of abnormal samples being misclassified as rare normal samples, and improves the accuracy and stability of detection.
Smart Images

Figure CN120724351B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of spatiotemporal trajectory data anomaly detection, and particularly relates to a trajectory anomaly route detection method and system based on self-supervised trajectory representation learning. BACKGROUND
[0002] Trajectory anomaly detection has become a key research topic with practical significance. In a real road environment, when a trajectory deviates significantly from its expected path, it is considered abnormal. For example, the normal route from the city center to the airport: most vehicles (such as trajectories A, B, and C) smoothly follow the main road of the urban expressway-airport expressway, and there is no detour or stop throughout the journey. However, a certain vehicle (trajectory D) suddenly turns into a secluded auxiliary road halfway, detours around an industrial area, and then returns to the main line. This significant deviation makes it an identifiable abnormal trajectory. Accurate detection of trajectory anomalies can help identify events at an early stage, thereby providing timely alerts for traffic management and personal navigation. Early self-supervised trajectory anomaly detection methods directly model normal patterns using raw GPS trajectories and identify anomalies. However, due to the inherent redundancy and noise in raw GPS data, subsequent methods introduce a grid-based representation. Specifically, the geographical space is divided into grids of a fixed size, and GPS points falling into the same grid are encoded as the same grid number, thereby improving robustness and computational efficiency. Based on this, recent research further enhances representation learning through recurrent neural networks (RNN) and autoencoder (AE) architectures, capturing sequence dependencies and optimizing reconstruction loss. Despite the significant progress made by existing research, there are still some problems.
[0003] Most existing trajectory anomaly detection methods rely solely on GPS or Grid single-view trajectory representation. However, the GPS trajectory view is highly sensitive to positioning errors and easily affected by noise. The Grid trajectory view has limited resolution and cannot capture fine-grained motion patterns and minor changes. This discretization strategy indeed makes trajectory features more robust to minor disturbances, but it inevitably causes information compression distortion and spatial neighbor disruption. When there is road congestion or vehicle stops, multiple consecutive points are compressed into the same grid, and the time sequence, speed difference, and start-stop details are all lost, making it impossible for the model to detect local anomalies. If two almost overlapping points happen to be on opposite sides of the grid boundary, they will be forced to be assigned different numbers, and the "zero distance" in the real geography is amplified to a "cross-grid difference" in the grid space. As a result, the insignificant jitter is misjudged as a significant deviation, and the real large detour may only span one or two grids and be ignored, making trajectory feature learning lack fine-grained description and blur the boundaries between anomalies and normality. Therefore, single grid division cannot simultaneously achieve global robustness and local fidelity.
[0004] And the current mainstream method uses an RNN-based model to capture the global spatio-temporal features of the entire trajectory. However, the trajectory anomaly in the real world often manifests as a local deviation, such as a sharp turn, a sudden stop, or a route violation, and these local anomalies do not significantly change the overall distribution of the entire trajectory. Relying only on global features may result in missing local anomalies and reducing detection accuracy.
[0005] In summary, the current spatio-temporal trajectory data anomaly detection model mainly has the following problems:
[0006] 1) Single trajectory view lacks interaction. Single trajectory view ignores the complementary information between different views, limiting the comprehensive understanding of the model to the trajectory behavior.
[0007] 2) Lack of detection capability for local anomalies. Existing methods focus on global trajectory representation and ignore abnormal patterns in local segments. SUMMARY
[0008] In order to solve the above-mentioned problems, the present application provides a trajectory anomaly route detection method and system based on self-supervised trajectory representation learning.
[0009] In a first aspect, the present application provides a trajectory anomaly route detection method based on self-supervised trajectory representation learning, which adopts the following technical solution:
[0010] A trajectory anomaly route detection method based on self-supervised trajectory representation learning, comprising:
[0011] Obtaining trajectory data;
[0012] Preprocessing the obtained trajectory data;
[0013] Building a trajectory dual-view contrast representation learning based on a dual-view synchronous mask strategy;
[0014] Encoding the temporal dynamics based on a spatio-temporal fusion mechanism and fusing with the spatial features;
[0015] Learning essential representations from the fused spatio-temporal features for trajectory reconstruction;
[0016] Error checking based on reconstructed trajectories to determine abnormal trajectories through reconstruction error.
[0017] Further, the preprocessing of the obtained trajectory data includes preprocessing each trajectory into two different representations, wherein the original trajectory is retained as a GPS sequence, denoted as A grid-based trajectory representation is constructed by dividing the geographic area into uniform grid cells and mapping each GPS point to its corresponding grid index , thereby obtaining a grid trajectory When multiple consecutive GPS points fall into the same grid cell, multiple identical grid entries are retained in the sequence, thereby generating the same trajectory for two complementary views, which is used to simultaneously capture fine-grained coordinate-level information and region-level semantics.
[0018] Furthermore, the trajectory dual-view contrast representation learning based on the dual-view synchronous masking strategy includes first calculating a shared binary mask matrix. , is used to indicate the segments in the trajectory that need to be masked, where n represents the trajectory length. The masked trajectory is constructed as follows: ;
[0019] By utilizing the spatial semantics contained in the original GPS track, the input sequence is encoded. To enhance the model's sensitivity to local anomalies, Each point in the matrix is mapped to an embedded representation through a linear transformation layer. Where n is the trajectory length, For the embedding dimension, position encoding is then added. To enhance the time information, the final input representation is obtained. : During the encoding process, the input embedding is multiplied by a valid mask to suppress invalid positions: To capture the temporal dependencies in variable-length trajectories, a GRU encoder is used to process the masked embedding. :
[0020] ,
[0021] in The final GPS embedding output of the GRU encoder is then used to obtain the global representation through average pooling: .
[0022] Furthermore, the trajectory dual-view comparison representation learning based on the dual-view synchronous masking strategy also includes a graph-structured grid trajectory embedding mechanism to encode spatial relationships. First, the grid trajectory sequence... After processing using a dual-view synchronous masking strategy, the mask sequence is obtained. The masked nodes are marked with a special mask. To replace this, in order to capture the spatial dependencies between grid cells, a grid-enhanced graph convolutional network (GCN) is used for representation learning. The GCN captures structural dependencies within the grid space and encodes the spatial context for each grid cell.
[0023] ,
[0024] Where L is the number of layers in the GCN. denotes the symmetrically normalized adjacency matrix, W is a learnable weight matrix, and the output of the last GCN layer is denoted as Z, i.e., the final grid embedding representation, where s is the total number of grids, and each position’s embedding vector is extracted from the grid representation matrix Z according to For the masked grid positions, the mask embedding is used for replacement, and the spatial embedding vector at position t in the trajectory is defined as: Finally, each trajectory is mapped to an embedding sequence: .
[0025] Further, the trajectory dual-view contrastive representation learning based on the dual-view synchronous mask strategy further includes, after obtaining the grid-based trajectory embedding sequence , further introducing position encoding and a Transformer encoder to capture the context dependency between spatial positions:
[0026] ,
[0027] where FC(·) denotes a fully connected layer, Transformer(·) denotes a Transformer encoder, is the final grid-based trajectory embedding, which retains the feature information of all positions on the trajectory, is used to support the mask reconstruction task to capture local trajectory features, and is used for contrastive learning with the GPS trajectory-level embedding. In this embodiment, the valid positions in are averaged-pooled to obtain the global representation of the grid trajectory : .
[0028] Further, the trajectory dual-view contrastive representation learning based on the dual-view synchronous mask strategy further includes, through a similarity-based contrastive learning strategy, fully integrating spatial information between the GPS view and the grid view, and realizing effective representation learning at multiple granularities. First, the dual-view trajectory global representations and are mapped to a shared embedding space:
[0029] ,
[0030] where (·) and (·) are two independent projection heads, which are usually implemented by a multi-layer perceptron (MLP) and both output a vector with a dimension of Further, the vector is ℓ2-normalized to calculate the cosine similarity:
[0031] ,
[0032] wherein represents the final trajectory representation in GPS view, represents the final trajectory representation in grid view, then the bidirectional similarity between GPS and grid representation is calculated:
[0033] ,
[0034] ,
[0035] wherein represents the similarity weight between a grid trajectory and all GPS trajectories in the dataset, represents the similarity weight between a GPS trajectory and all grid trajectories, finally the most dissimilar trajectories are selected from the opposite view to construct two negative sample pairs: , is a temperature scaling factor.
[0036] Further, the spatio-temporal fusion mechanism encodes the time dynamics and fuses with the spatial features, including a learnable time embedding mechanism, which shows the encoding of the time information associated with each trajectory point, the original time vector is processed by a double-branch neural network structure to obtain the time embedding :
[0037] ,
[0038] wherein and are the learnable weight matrices of the linear and nonlinear branches, respectively, and the final embedding dimension is The spatio-temporal interaction attention mechanism is then used to fuse the spatial and temporal features extracted from the grid trajectory and time embedding modules, respectively, wherein the spatial features are obtained from the dual-view contrastive representation learning module , and the temporal feature is represented as The spatial and temporal features are projected into query Q, key K and value V vectors through three sets of shared linear transformations, respectively, and finally generate spatial attention output and temporal attention output , respectively, through feed-forward network FFN and residual connection to enhance and stabilize the features, and the final spatio-temporal fusion feature is obtained by concatenation output:
[0039] ,
[0040] wherein and represent normalization operations, represents the spatio-temporal embedding of the i-th point in the trajectory, and denote the spatial and temporal attention outputs corresponding to the i-th trajectory point, respectively.
[0041] Further, the learning of the intrinsic representation from the fused spatio-temporal features for trajectory reconstruction includes, at each time step , recursively updating the hidden state with the spatio-temporal embedding as input. The final hidden state captures the sequence dependency and serves as a compact latent representation of the trajectory for the reconstruction task: where denotes the hidden state at the previous time step, and a GMM is adopted to model the latent space of the trajectory, then a path class variable C is introduced to capture different spatio-temporal trajectory patterns, where the joint posterior of the latent vector of a given trajectory and the class is decomposed as: where is the probability that the trajectory belongs to the class , the trajectory , the latent vector , is , which is calculated by Bayes' theorem as:
[0042] ,
[0043] where defines the class prior of the path type, and the latent representation is determined by the mixture components and the class prior , and finally a generative network is introduced, where is a hyper-parameter, and for a given sequence and latent path vector , the hidden state is recursively updated as: where denotes the ST-Decoder implemented by a learnable RNN, and its initial hidden state is set to the latent vector, i.e. .
[0044] Further, the error checking based on the reconstructed trajectory judges the abnormal trajectory through reconstruction error, including inputting the trajectory data into a model, the model extracts embedded representations of GPS and Grid modalities through a dual-view encoder respectively, and gives a reconstruction output of the trajectory in the latent space; calculating the reconstruction error to measure the deviation degree between the input trajectory and its reconstruction version, and based on the Gaussian mixture prior, assigning the latent vector of the trajectory to the Gaussian component with the maximum posterior probability, and calculating the negative logarithmic probability density between the component center, and comprehensively obtaining a comprehensive abnormal score, if the score exceeds an adaptive threshold, the trajectory is determined to be abnormal, otherwise it is considered to be normal.
[0045] In a second aspect, a trajectory anomaly route detection system based on self-supervised trajectory representation learning includes:
[0046] A data acquisition module is configured to acquire trajectory data.
[0047] A preprocessing module is configured to preprocess the acquired trajectory data.
[0048] A learning module is configured to construct trajectory dual-view contrast representation learning based on a dual-view synchronous mask strategy.
[0049] A fusion module is configured to encode time dynamics based on a spatiotemporal fusion mechanism and fuse with spatial features.
[0050] A reconstruction module is configured to learn essential representation from fused spatiotemporal features for trajectory reconstruction.
[0051] A judgment module is configured to perform error checking based on the reconstructed trajectory to judge the abnormal trajectory through the reconstruction error.
[0052] In a third aspect, the present application provides a computer readable storage medium, which stores a plurality of instructions, the instructions being suitable for being loaded and executed by a processor of a terminal device.
[0053] In a fourth aspect, the present application provides a terminal device including a processor and a computer readable storage medium, the processor being used to implement instructions; the computer readable storage medium is used to store a plurality of instructions, the instructions being suitable for being loaded and executed by the processor to perform a trajectory anomaly route detection method based on self-supervised trajectory representation learning.
[0054] In summary, the present application has the following beneficial technical effects:
[0055] The application proposes a brand-new trajectory anomaly detection model. The model enriches the trajectory representation by fusing GPS trajectory and grid-based trajectory features; at the same time, a "dual-view synchronous mask" mechanism is designed, which enables the model to perceive local disturbances in the spatial and temporal dimensions simultaneously during the training phase, thereby improving the sensitivity to local anomalies; in addition, a generative framework based on latent Gaussian mixture distribution is introduced to robustly model complex data distribution and provide support for accurate anomaly detection. The core advantages are:
[0056] Richer representation: through GPS+Grid dual-view fusion, both fine coordinates and structured road information are considered, significantly enhancing the trajectory feature expression ability.
[0057] More sensitive to anomalies: the dual-view synchronous mask forces the model to simultaneously reconstruct the masked spatial and temporal segments during training, keeping the model highly sensitive to local anomalies.
[0058] More robust distribution: the latent Gaussian mixture prior can capture multiple normal patterns, reducing the risk of misjudging abnormal samples as "rare normal", thereby improving detection accuracy and stability. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 The overall flowchart of the application.
[0060] Figure 2 The data preprocessing module flowchart of step 1 of the application.
[0061] Figure 3 The trajectory anomaly detection model training flowchart proposed by the application.
[0062] Figure 4 The trajectory dual-view contrast learning module flowchart of step 2 of the application.
[0063] Figure 5 The spatio-temporal fusion module flowchart of step 3 of the application.
[0064] Figure 6 The trajectory reconstruction module flowchart of step 4 of the application.
[0065] Figure 7 The comparison bar chart between the results corresponding to the ablation experiments on the porto dataset of the application.
[0066] Figure 8 The comparison bar chart between the results corresponding to the ablation experiments on the dataset of a certain city of the application. DETAILED DESCRIPTION
[0067] The application will be further described in detail below in conjunction with the drawings.
[0068] Example 1
[0069] Referring Figure 1 , the trajectory anomaly route detection method based on self-supervised trajectory representation learning of the embodiment includes:
[0070] Obtaining trajectory data;
[0071] Preprocessing the obtained trajectory data;
[0072] Constructing trajectory dual-view contrastive representation learning based on a dual-view synchronous mask strategy;
[0073] Encoding the temporal dynamics based on a spatiotemporal fusion mechanism and fusing with the spatial features;
[0074] Learning essential representations from the fused spatiotemporal features for trajectory reconstruction;
[0075] Performing error checking based on the reconstructed trajectory, and determining the abnormal trajectory through reconstruction error.
[0076] Specifically,
[0077] S1 trajectory data preprocessing.
[0078] To facilitate the dual-view contrastive learning between the GPS view and the grid view, the embodiment preprocesses each trajectory into two different representation forms. The original trajectory is kept as a GPS sequence, denoted as . According to previous research, the embodiment divides the geographical area into uniform grid cells and maps each GPS point to its corresponding grid index , constructing a grid-based trajectory representation, thereby obtaining the grid trajectory . When multiple consecutive GPS points fall into the same grid cell, the embodiment keeps multiple identical grid entries in the sequence. This process produces the same trajectory in two complementary views, enabling the model to capture both fine-grained coordinate-level information and regional-level semantics.
[0079] This dual representation form enables the model to focus on both details and global information during the learning process, forming complementary perspectives. Specifically, the GPS sequence can provide real-time location information, capturing the dynamic changes and details of the motion trajectory; while the grid-based representation effectively simplifies the information, extracting features representative of the region. This complementary relationship enables the model to more comprehensively understand and predict the behavior patterns of the trajectory when performing contrastive learning. In this way, the embodiment not only enhances the richness and accuracy of trajectory representation, but also provides a more robust foundation for subsequent learning tasks. By combining the advantages of the two perspectives, the model of the embodiment can better capture and understand the internal structure and patterns of trajectories in complex geographical environments.
[0080] S2 Trajectory dual-view pair-wise contrastive representation learning
[0081] The present application aims to simultaneously capture fine-grained location information features from the GPS view and obtain large-scale regional-level features from the grid view by adopting a dual processing flow technical architecture. The main purpose of this dual-view design is to make full use of two different types of spatial information in order to improve the overall performance and robustness of the model in multiple dimensions. Specifically, the GPS view provides continuous location information, capturing subtle displacement changes and reflecting dynamic movement trajectories and behavior patterns. This fine-grained data can reveal the precise movement path of individuals at a specific time, providing high-resolution spatial analysis. This is particularly important for real-time monitoring, anomaly detection, and other application scenarios, as it can timely identify potential deviations and abnormal phenomena. At the same time, the grid view extracts regional-level features by dividing the geographical area into uniform grid cells, reflecting the overall characteristics and trends within the region, which helps to identify more extensive patterns and behaviors. By combining regional-level features, the model can obtain richer contextual information, effectively enhancing its understanding of complex environments. First, a dual-view synchronous masking strategy is adopted to enhance the sensitivity of local anomalies. This strategy applies synchronous masking to the corresponding parts of the two views, making the model more targeted and real-time when capturing anomalies. This not only helps to improve detection accuracy, but also effectively reduces the impact of noise on the model.
[0082] S2.1 Dual-view synchronous masking strategy
[0083] In order to achieve effective alignment between the GPS view and the grid (Grid) view during the training process and ensure fair comparison, the present embodiment designs a dual-view synchronous masking strategy. Specifically, a shared binary mask matrix is first calculated , which indicates which segments in the trajectory need to be masked, where n represents the length of the trajectory. This will be applied to the same position and segment length in both views, ensuring that the two representations are completely consistent in masking. The masked positions are replaced by the corresponding mask markers in each view, thereby maintaining semantic alignment and consistency between the two representations. The construction of the mask trajectory is shown below, where represents the Hadamard product:
[0084] (1)
[0085] S2.2 GPS trajectory embedding
[0086] The present embodiment aims to utilize the fine spatial semantics contained in the original GPS trajectory and encode the input sequence to enhance the model's sensitivity to local anomalies.
[0087] First, each point in the trajectory is mapped to an embedding representation by a linear transformation layer, where n is the length of the trajectory, and d is the embedding dimension. Subsequently, positional encoding is added to reinforce the temporal information, resulting in the final input representation :
[0088] (2)
[0089] During the encoding process, the input embedding is multiplied by a valid mask to suppress invalid positions:
[0090] (3)
[0091] To capture the temporal dependencies in variable-length trajectories, this embodiment uses a GRU encoder to process the masked embedding :
[0092] (4)
[0093] where is the final GPS embedding output of the GRU encoder. Subsequently, the global representation is obtained by average pooling:
[0094] (5)
[0095] where summarizes the macroscopic motion patterns of the entire trajectory and serves as the global embedding for contrastive learning. In addition, the GPS embedding without pooling retains the encoding output at each time step, facilitating subsequent masked reconstruction tasks and enabling the model to learn fine-grained local dynamics.
[0096] Next, each view is separately encoded and pooled to generate compact embeddings that preserve spatial structure. This strategy of separate encoding allows the model to focus on different levels of information more flexibly. In the GPS view, fine-grained location information can reflect dynamic changes, while in the grid view, regional-level features can provide important contextual information. Finally, a similarity-based contrastive learning mechanism aligns the GPS and grid representations, promoting the integration of mutual information between views. This enables the two views to share information during learning, enhancing the model's learning ability and generalization performance. By integrating the advantages of both perspectives, this method can more comprehensively understand trajectory data, thereby improving the recognition and prediction of motion behaviors in complex environments.
[0097] The module is committed to extracting high-quality continuous spatial representations from GPS format trajectories by modeling the original coordinate sequence to capture the dynamic evolution pattern of the trajectory in real geographical space. This process not only reveals the fine-grained motion characteristics of the trajectory at the micro level, but also lays a solid embedding foundation for subsequent tasks of spatio-temporal alignment and cross-modal representation learning. Specifically, in tasks such as contrastive learning and trajectory reconstruction, the representations output by the module can be used to build a fine-grained time series contrastive framework and a missing point prediction model, thereby enhancing the model's ability to capture the semantic consistency and spatial integrity of the trajectory. Through a joint training mechanism, these representations can be complementarily fused with grid embeddings to support a more robust and more generalizable multi-modal trajectory understanding system.
[0098] S2.3Grid trajectory embedding
[0099] To depict the spatial structure features of grid-based trajectories at a coarse-grained level, it is necessary to capture the spatial dependency between position points. To this end, the embodiment designs a grid trajectory embedding module of graph structure to effectively encode such spatial relationships.
[0100] First, the grid trajectory sequence is processed by a dual-view synchronous mask strategy to obtain a mask sequence , where the masked nodes are marked by a special mask . Secondly, to better capture the spatial dependency between grid cells, the embodiment uses a grid-enhanced graph convolutional network (GCN) for representation learning. GCN captures structural dependencies within the grid space and encodes spatial context for each grid cell:
[0101] (6)
[0102] where L is the number of layers of GCN, is a symmetric normalized adjacency matrix, and W is a learnable weight matrix. Initially, Z(0) is composed of an input feature matrix of grid feature vectors. For brevity, the output of the last layer of GCN is denoted as Z, i.e., the final grid embedding representation, where s is the total number of grids.
[0103] Next, the embodiment extracts the embedding vector of each position from the grid representation matrix Z according to the grid index of ; for the masked grid positions, the mask embedding is used for replacement. The spatial embedding vector at position t in the trajectory is defined as:
[0104] (7)
[0105] Finally, each trajectory is mapped to an embedding sequence In obtaining the grid-based trajectory embedding sequence After that, the model further introduces position encoding and a Transformer encoder to capture the contextual dependency between spatial positions:
[0106] (8)
[0107] where FC(·) denotes a fully connected layer and Transformer(·) denotes a Transformer encoder. For the final grid-based trajectory embedding, the feature information of all positions on the trajectory is preserved, which can be used to support the mask reconstruction task to capture local trajectory features. To compare with the GPS trajectory-level embedding, the embodiment performs average pooling on the effective positions in to obtain the global representation of the grid trajectory :
[0108] (9)
[0109] S2.4 Similarity-based contrastive learning
[0110] Through the similarity-based contrastive learning strategy, the spatial information is fully integrated between the GPS view and the grid view, and effective representation learning is achieved on multiple granular levels.
[0111] First, the global representations of the trajectories in the two views and are mapped to a shared embedding space:
[0112] (10)
[0113] where (·) and (·) are two independent projection heads, usually implemented by a multi-layer perceptron (MLP), both outputting a vector of dimension The embodiment further performs ℓ2 normalization on these vectors in order to calculate the cosine similarity:
[0114] (11)
[0115] where denotes the final trajectory representation under the GPS view, denotes the final trajectory representation under the grid view. These two embeddings constitute a positive sample pair, denoted as Pos=[ , ]. Subsequently, the embodiment calculates the bidirectional similarity between the GPS and grid representations:
[0116] ,
[0117] (12)
[0118] where denotes the similarity weight between a grid trajectory and all GPS trajectories in the dataset, denotes the similarity weight between a GPS trajectory and all grid trajectories. Together, they form a bidirectional similarity matrix, which is used to select informative negative samples. Based on this matrix, we select the most dissimilar trajectories from the opposite view to construct two negative sample pairs: . is the temperature scaling factor.
[0119] By minimizing the distance between positive samples and maximizing the distance between negative samples, the model effectively fuses the features learned from GPS and grid views, encouraging the representations of the same trajectory in different views to be close to each other, thus achieving robust dual-view feature learning.
[0120] S3 Spatio-temporal feature fusion
[0121] Temporal information is crucial in trajectory anomaly detection, as it helps the model capture motion trends and identify patterns such as long stays or congestion. To better capture complex motion patterns, the invention proposes a spatio-temporal fusion mechanism to encode temporal dynamics and fuse them with spatial features, thus achieving bidirectional dependency modeling.
[0122] S3.1 Temporal embedding
[0123] This embodiment draws on existing temporal feature representation methods and designs a learnable temporal embedding module to explicitly encode the temporal information associated with each trajectory point. Each temporal input is represented as a 6-dimensional vector, containing the following discrete temporal attributes: hour , day of the week , week of the year , month , holiday flag , and weekend flag
[0124] : n. (13)
[0125] The original temporal vector is processed through a dual-branch neural network structure to obtain the temporal embedding :
[0126] (14)
[0127] where and These are the learnable weight matrices for the linear and nonlinear branches, respectively, with the final embedding dimension being... .
[0128] S3.2 Spatiotemporal Feature Fusion
[0129] To fully model the interaction between spatial and temporal features, this embodiment designs a spatiotemporal interaction attention mechanism to fuse spatial and temporal features extracted from the grid trajectory and temporal embedding modules, respectively. First, spatial features are obtained from the dual-view contrastive representation learning module. .here, It integrates the rich spatial granularity of the GPS perspective while retaining the structural advantages of grid representation. In contrast, although GPS features While it contains richer spatial details, it also introduces excessive noise, which may make the model overly sensitive, weakening its ability to accurately identify trajectory anomalies and thus reducing its stability and generalization ability. Temporal features are obtained from the temporal embedding module and are represented as... .
[0130] Spatial and temporal features are projected into query (Q), key (K), and value (V) vectors through three shared linear transformations, as follows:
[0131] (15)
[0132] (16)
[0133] in , and These are learnable weights used to project temporal features into a vector of queries, keys, and values to compute their interaction with spatial features; , and Then a similar projection is performed to capture the effect of spatial features on temporal features.
[0134] Next, this embodiment generates spatial attention outputs respectively. and time attention output :
[0135] ,
[0136] (17)
[0137] Subsequently, the features are enhanced and stabilized using a feedforward network (FFN) and residual connections, respectively. The final spatiotemporal fusion features are obtained by concatenating the outputs.
[0138] (18)
[0139] where and denote the normalization operation, denotes the spatio-temporal embedding of the i-th point in the trajectory. Moreover, and denote the spatial and temporal attention outputs corresponding to the i-th trajectory point, respectively.
[0140] S4 Trajectory Reconstruction
[0141] To further learn the essential representation from the fused spatio-temporal features, this embodiment adopts an encoder-decoder architecture to construct the latent representation space. The invention further uses a Gaussian Mixture Model (GMM) to model the distribution of various normal trajectories, thereby capturing various motion patterns commonly seen in real-world scenarios.
[0142] S4.1 Latent Space Generation
[0143] At each time step , the ST-Encoder recursively updates the hidden state with the spatio-temporal embedding as input. The final hidden state captures the sequence dependency and is used as the compact latent representation of the trajectory for the reconstruction task:
[0144] (19)
[0145] where denotes the hidden state of the previous time step.
[0146] S4.2 Gaussian Mixture Latent Space Modeling
[0147] Considering that trajectories usually exhibit multiple spatio-temporal patterns, a single distribution is difficult to fully characterize. This embodiment uses a GMM to model the latent space of the trajectory:
[0148] (20)
[0149] (21)
[0150] where denote the spatial and temporal latent vectors, is the inference network, is the hidden state, and are nonlinear mapping functions. I is the identity matrix, and Tr denotes the input trajectory sequence.
[0151] Due to the diversity of travel areas, road types, and time, this embodiment introduces a path category variable C to capture different spatio-temporal trajectory patterns. For each category, the latent trajectory vector is modeled by a Gaussian distribution: To jointly model the latent representation and its belonging category, this embodiment decomposes the joint posterior of the latent vector and the category as:
[0152] (22)
[0153] where can be interpreted as the probability that the trajectory belongs to the category . Since the trajectory can be encoded as the latent vector , the can be equivalently written as and can be computed by Bayes' rule: (23)
[0154] where defines the category prior of the path type. Based on this, the latent representation is jointly determined by the mixture component and the category prior .
[0155] S4.3 Trajectory Reconstruction
[0156] To realize trajectory generation and detection, this embodiment introduces a generative network where is a hyper-parameter.
[0157] Given the sequence and the latent path vector , the hidden state is recursively updated as:
[0158] (24)
[0159] where denotes the ST-Decoder implemented by a learnable RNN, whose initial hidden state is set to the latent vector, i.e., . At each step i, the embedding is generated by the following conditional distribution:
[0160] (25)
[0161] This method ensures that each embedding is based on all previously generated embeddings and the latent vector is generated, so that the trajectory generation process can take advantage of the time dependency.
[0162] S5 model iteration training and saving model.
[0163] In the first stage, the original trajectory data is first completed standardization preprocessing. Subsequently, the embodiment will be synchronized after the mask processing GPS trajectory embedding and Grid trajectory embedding Respectively input each multilayer perception (MLP), through the calculation mask reconstruction loss (MaskReconstructionLoss) to two view encoders independent pre-training, make them can from the masked input to recover complete trajectory representation.
[0164] In the second stage, the GPS trajectory level embedding and Grid trajectory level embedding Be configured as positive sample pair (two views of the same trajectory) and negative sample pair (any view of different trajectories), through the contrastive loss (ContrastiveLoss) of the form of InfoNCE Joint optimization, so that the dual view information is aligned and fused, obtain consistent and complementary spatiotemporal representation.
[0165] Finally, in the third stage, the fused embedding is sent to the trajectory reconstruction module, and the trajectory feature is further refined by minimizing the trajectory reconstruction loss (TrajectoryReconstructionLoss) ; At the same time, the Gaussian loss (GaussianLoss) is calculated by using the Gaussian mixture prior, and the distribution of the latent space is regularized, so as to provide discriminable, interpretable and robust latent representation for subsequent anomaly detection.
[0166] S5 uses model to predict future data.
[0167] When using the model for anomaly detection, the embodiment inputs the trajectory data into the model, and the model extracts the embedding representation of the GPS and Grid modalities through the dual view encoder, and gives the reconstruction output of the trajectory in the latent space; The reconstruction error is calculated to measure the deviation between the input trajectory and its reconstruction version. And with the learned Gaussian mixture prior, the latent vector of the trajectory is assigned to the Gaussian component with the maximum posterior probability, and the negative log probability density between it and the component center is calculated, and the comprehensive anomaly score is obtained. If the score exceeds the adaptive threshold, the trajectory is determined to be abnormal, otherwise it is considered to be normal.
[0168] Among them, in order to detect abnormal trajectories, the embodiment adopts a generative model , used to evaluate the reconstruction trajectory from the latent representation The path type is determined by computing the likelihood of the trajectory under all possible path types and selecting the type with the maximum likelihood. Where argmax denotes the path type c that maximizes the likelihood function. Here, n represents the length of the trajectory, and as a normalization factor, it ensures a fair comparison between trajectories of different lengths. represents the mean vector of the latent vector for each path type c.
[0169] For online anomaly detection, at time step , the anomaly score is computed based on the spatio-temporal embedding of the historical trajectory segments and their likelihood under each path type. The anomaly score of a trajectory is defined as:
[0170]
[0170]
[0171] In this case, the higher the anomaly score, the more likely the trajectory is to be abnormal under the learned generative model.
[0172]
Experimental Verification
[0173] In the field of trajectory anomaly detection, the public real abnormal label is scarce, and existing research generally relies on manual annotation, which is time-consuming and lacks standardization, and is difficult to reproduce. Therefore, the invention uses a set of rule synthesis strategies to automatically generate abnormal samples to build a unified and controllable experimental benchmark. The specific scheme is as follows:
[0174] Two adjustable parameters are introduced:
[0175] α——The proportion of continuous abnormal segments in the whole trajectory;
[0176] d——The displacement of the segment in the spatial grid (unit: grid unit).
[0177] Unlike existing methods that only perturb a single point, the invention applies a bias to the entire continuous trajectory segment, which can more realistically simulate real traffic anomaly scenarios such as vehicle detours and congestion rerouting. For example, when α=0.2, d=2, it means that the entire continuous segment of 20% of the trajectory is translated by 2 grid units. Considering that anomalies often accompany temporal anomalies, the invention adjusts the time interval between corresponding points according to the original trajectory speed consistency principle while perturbing the spatial coordinates, ensuring that the number of trajectory points remains unchanged and the speed distribution is reasonable. In all experiments, the proportion of synthetic abnormal samples in the total data volume is uniformly set to 5%, which avoids class imbalance and ensures the repeatability of the experiment.
[0178] Experimental Effect Comparison
[0179] Tables 1 and 2 give the experimental results of comparing other trajectory anomaly detection methods on the Porto and City data sets. The Precision-Recall curve area (PR-AUC) is used as the evaluation index of the anomaly detection performance in this embodiment. Due to the extremely small proportion of abnormal samples, the data set is highly imbalanced, and PR-AUC can reliably measure the ability of the model to identify anomalies in this scenario. To comprehensively evaluate the detection performance, this embodiment changes the proportion of observed trajectories ρ, which controls the proportion of visible trajectory data in the detection stage. Under all settings, learning-based methods are significantly better than distance-based baselines (iBAT and TPRRO). Among them, the present application always achieves the best results, and is superior to existing methods. Compared with the current SOTA method MST-OATD method, the PR-AUC of the present application on the Porto data set is improved by up to 20%, and the average improvement on the City data set is 5%.
[0180] For example, on the Porto data set ρ=0.5, the present application significantly outperforms all baselines. This improvement is mainly due to the dual-view synchronous mask pre-training strategy proposed in this embodiment, which enables the model to accurately capture local anomalies even with limited observed data. In addition, in the more challenging scenario (α=0.3, d=2), the present application still achieves a PR-AUC of 0.987, further verifying its robustness. On the City data set, although the MST-OATD method performs strongly, and the data size is large and the time irregularity is high, the present application still leads overall, achieving a PR-AUC of 0.990 under the same settings as above. As ρ increases, PR-AUC improves accordingly, indicating that observing longer trajectory segments helps improve detection accuracy. Even in the low anomaly setting (d=3, α=0.1), when ρ=1.0, the present application still achieves a PR-AUC of 0.878 on Porto, demonstrating its stability. This robust performance is due to the complementary fusion of fine-grained GPS features and coarse-grained Grid representation: the Grid view effectively suppresses false positives caused by normal changes such as lane changes, while the GPS view maintains high sensitivity to sudden anomalies. Compared with single-view methods, this dual-view design significantly improves anomaly detection rates.
[0181] Table 1: Performance comparison under different parameter settings on the Porto data set
[0182]
[0183] Table 2: Performance comparison under different parameter settings on the City data set
[0184]
[0185]
Ablation experiment
[0186] To evaluate the effectiveness of each module of the present application, the present embodiment conducts a systematic ablation experiment on the Porto and City data sets. Specifically, the present embodiment compares the present application with the following six variants:
[0187] w / oGPS module: Remove GPS data in the training stage, only keep the grid-based representation.
[0188] w / oGrid module: Remove grid data in the training stage, only keep the original GPS data.
[0189] w / oMLM module: Remove the pre-training of the synchronization mask on GPS and grid data.
[0190] w / oCL module: Remove the cross-modal dual-view contrastive loss.
[0191] w / oSM module: Remove the similarity matrix and use randomly sampled negative samples instead.
[0192] w / oTime module: Remove all time information and only use spatial information.
[0193] All ablation experiments are conducted under the parameter settings of α = 0.2 and d = 2, and the PR-AUC of each variant under ρ = 0.5, ρ = 0.7 and ρ = 1.0 is recorded, and the results are shown in the following figures.
[0194] As shown in the following figure, the figure shows the PR-AUC value comparison of each variant in the Porto data set under ρ = 0.7. Figure 7 Figure 8 The figure shows the PR-AUC value comparison of each variant in the City data set under ρ = 0.7. Figure 7 As can be seen from the above figures, the w / oMLM has a smaller decline under ρ = 0.7. This shows that mask pre-training can help the model capture local anomalies under low observation ratio, thereby maintaining robustness. Removing the time input (w / oTime module) causes a performance drop of about 18%, indicating that the time clue is crucial for depicting the trajectory dynamics. Without the temporal context, the model cannot effectively identify time-dependent anomalies such as irregular speed or abnormal stay duration. The w / oGrid module causes a performance degradation of about 20%, indicating that the grid representation provides key coarse-grained spatial context, which can suppress the interference of noise and minor deviations. After removing the grid view, the model relies too much on the fine-grained GPS signal and has difficulty establishing stable spatial patterns. It is worth noting that the w / oGPS module variant, which only retains GPS and removes the grid, performs worse than the w / oGrid module. Although GPS has high accuracy, it is too sensitive to subtle disturbances (such as lane changes or positioning noise) and is prone to false positives. This further verifies the necessity of retaining the grid-level trajectory features obtained by GPS fusion in dual-view contrastive learning, which helps to stabilize the spatial semantics and improve detection accuracy.
[0195] In addition, the performance of the w / oCL module and the w / oSM module also decreases, but the magnitude is relatively small. Among them, removing the similarity-based negative sampling (w / oSM module) still brings certain degradation, indicating that the similarity-driven negative sample selection has a positive contribution to the model performance.
[0196] Embodiment 2
[0197] The embodiment provides a trajectory anomaly route detection system based on self-supervised trajectory feature learning, comprising:
[0198] A data acquisition module configured to acquire trajectory data;
[0199] A preprocessing module configured to preprocess the acquired trajectory data;
[0200] A learning module configured to construct trajectory dual-view contrast representation learning based on a dual-view synchronous mask strategy;
[0201] A fusion module configured to encode time dynamics based on a spatiotemporal fusion mechanism and fuse with spatial features;
[0202] A reconstruction module configured to learn essential features from the fused spatiotemporal features for trajectory reconstruction;
[0203] A judgment module configured to perform error inspection based on the reconstructed trajectory, and judge the abnormal trajectory through the reconstruction error.
[0204] A computer-readable storage medium, wherein a plurality of instructions are stored, the instructions being adapted to be loaded and executed by a processor of a terminal device, and the instructions being adapted to implement a trajectory anomaly route detection method based on self-supervised trajectory feature learning.
[0205] A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement instructions, and the computer-readable storage medium being configured to store a plurality of instructions, the instructions being adapted to be loaded and executed by the processor, and the instructions being adapted to implement a trajectory anomaly route detection method based on self-supervised trajectory feature learning.
[0206] The above are preferred embodiments of the present application, and are not intended to limit the protection scope of the present application, therefore: any equivalent changes made on the basis of the structure, shape, principle of the present application should be covered within the protection scope of the present application.
Claims
1. A trajectory anomaly route detection method based on self-supervised trajectory representation learning, characterized in that, The method comprises the following steps: acquiring trajectory data; preprocessing the acquired trajectory data; constructing trajectory dual-view contrastive representation learning based on a dual-view synchronous mask strategy; encoding time dynamics based on a spatiotemporal fusion mechanism and fusing with spatial features; learning essential representations from the fused spatiotemporal features for trajectory reconstruction; performing error inspection based on the reconstructed trajectory, and determining abnormal trajectories through reconstruction error; The trajectory dual-view contrastive representation learning based on the dual-view synchronous mask strategy includes first calculating a shared binary mask matrix for indicating segments in the trajectory that need to be masked, wherein n represents the length of the trajectory, and the mask trajectory is constructed in the following manner: ; By encoding the input sequence with the spatial semantics contained in the original GPS trajectories to enhance the model's sensitivity to local anomalies, each point in is mapped to an embedding representation by a linear transformation layer where n is the trajectory length, is the embedding dimension, subsequently, position encoding is added to reinforce the temporal information, resulting in the final input representation : During the encoding process, the input embedding is multiplied by a valid mask to suppress invalid positions: Meanwhile, to capture the temporal dependencies in variable-length trajectories, a GRU encoder is used to process the masked embedding : , wherein The final GPS embedding output for the GRU encoder, followed by average pooling to get the global representation: ; The learning of the essential representation from the fused spatio-temporal features for trajectory reconstruction includes, at each time step with the spatio-temporal embedding as input, recursively updating a hidden state The final hidden state captures the sequence dependency and is used as a compact latent representation of the trajectory for the reconstruction task: , where denotes the hidden state of the previous time step and models the latent space of trajectories with a GMM, then introduces a path class variable C to capture different spatio-temporal trajectory patterns, where the latent vector of a given trajectory is decomposed into the joint posterior of the class , where is the probability that the trajectory belongs to the class is the trajectory is the latent vector , i.e. is calculated by Bayes' theorem as: , wherein, a class prior defining the type of path, a latent representation from the mixture components with the class prior co-determine, finally introducing the generative network wherein are hyperparameters, for a given sequence and a latent path vector , a hidden state is recursively updated as: , wherein represents an ST-Decoder implemented by a learnable RNN, whose initial hidden state is set to the latent vector, i.e. .
2. The trajectory anomaly route detection method based on self-supervised trajectory representation learning according to claim 1, wherein, The pre-processing of the acquired trajectory data includes pre-processing each trajectory into two different representations, wherein the original trajectory is retained as a GPS sequence, denoted as , and a grid-based trajectory representation is constructed by dividing the geographical area into uniform grid cells and mapping each GPS point to its corresponding grid index , thereby obtaining a grid trajectory ; when multiple consecutive GPS points fall into the same grid cell, multiple identical grid entries are retained in the sequence, thereby generating the same trajectory in two complementary views for simultaneously capturing fine-grained coordinate-level information and area-level semantics.
3. The trajectory anomaly route detection method based on self-supervised trajectory representation learning according to claim 2, wherein, The trajectory dual-view contrastive representation learning based on the dual-view synchronous mask strategy further includes a grid trajectory embedding mechanism based on a graph structure to encode spatial relationships. First, the grid trajectory sequence After processing by the dual-view synchronous mask strategy, a mask sequence is obtained where the masked nodes are marked by a special mask Instead, to capture spatial dependencies between grid cells, a grid-enhanced graph convolutional network (GCN) is used for representation learning. GCN captures structural dependencies within the grid space and encodes spatial context for each grid cell: , where L is the number of layers of GCN, denotes the symmetrically normalized adjacency matrix, W is a learnable weight matrix, and the output of the last layer of GCN is denoted as Z, i.e., the final grid embedding representation, where s is the total number of grids, and then each position’s embedding vector is extracted from the grid representation matrix Z according to the grid index of the grid position is masked, the masked embedding is used to replace it, and the spatial embedding vector at position t in the trajectory is defined as: , Ultimately, each trajectory is mapped as an embedding sequence .
4. The trajectory anomaly route detection method based on self-supervised trajectory representation learning according to claim 3, wherein, The trajectory dual-view contrastive representation learning based on the dual-view synchronous mask strategy further includes obtaining a grid-based trajectory embedding sequence After that, position encoding is further introduced and a Transformer encoder to capture the context dependency between spatial positions: , where FC(·) denotes a fully connected layer, Transformer(·) denotes a Transformer encoder, For the final grid-based trajectory embedding, the feature information of all positions on the trajectory is reserved to support the mask reconstruction task to capture local trajectory features, and to learn from the GPS trajectory-level embedding for contrastive learning, the effective positions in the grid trajectory are averaged-pooled to obtain the global representation of the grid trajectory : . 5. The trajectory anomaly route detection method based on self-supervised trajectory representation learning of claim 4, wherein, The trajectory dual-view comparative representation learning based on the dual-view synchronous masking strategy also includes a similarity-based comparative learning strategy to fully integrate spatial information between the GPS view and the grid view, and to achieve effective representation learning at multiple granular levels. First, the global representation of the trajectory from both views is... and Mapping to shared embedded space: , where (·) are two independent projection heads implemented by multi-layer perceptron MLP, both outputting vectors of dimension further ℓ2-normalizing the vectors to compute cosine similarity: where denotes the final trajectory representation in GPS view, denotes the final trajectory representation in grid view, then the bidirectional similarity between GPS and grid representation is computed: , , wherein represents the similarity weight between a grid trajectory and all GPS trajectories in the dataset, represents the similarity weight between a GPS trajectory and all grid trajectories, and finally the least similar trajectories are selected from the opposite view to construct two negative sample pairs: , T is a temperature scaling factor.
6. The trajectory anomaly route detection method based on self-supervised trajectory representation learning according to claim 5, wherein, The spatio-temporal fusion mechanism encodes time dynamics and fuses with spatial features, including a learnable temporal embedding mechanism that shows the encoding of time information associated with each trajectory point, the original time vector is processed by a dual-branch neural network structure to obtain a time embedding : , wherein and are learnable weight matrices for linear and non-linear branches, respectively, with final embedding dimension , and spatio-temporal interaction attention mechanism is used to fuse the spatial and temporal features extracted from the grid trajectory and time embedding modules, respectively, wherein the spatial feature is obtained from the dual-view contrastive representation learning module , and the temporal feature is represented as , the spatial and temporal features are projected into query Q, key K, and value V vectors through three sets of shared linear transformations, respectively, and finally the spatial attention output and the temporal attention output are generated, respectively, and the features are enhanced and stabilized through a feed-forward network FFN and a residual connection, respectively, and the final spatio-temporal fusion feature is obtained through concatenation output: , wherein and denotes a normalization operation, denotes the spatio-temporal embedding of the i-th point in the trajectory, and denote the spatial and temporal attention outputs corresponding to the i-th trajectory point, respectively.
7. The trajectory anomaly route detection method based on self-supervised trajectory representation learning of claim 6, wherein, the error inspection based on the reconstructed trajectory, and determining abnormal trajectories through reconstruction error, comprises inputting the trajectory data into a model, the model extracting embedded representations of GPS and Grid modalities through a dual-view encoder respectively, and giving a reconstruction output of the trajectory in a latent space; calculating reconstruction error to measure the deviation degree between the input trajectory and its reconstruction version, simultaneously assigning the latent vector of the trajectory to a Gaussian component with the maximum posterior probability based on a Gaussian mixture prior, and calculating the negative logarithmic probability density between the latent vector and the center of the component, comprehensively obtaining a comprehensive abnormal score, if the score exceeds an adaptive threshold, determining that the trajectory is abnormal, otherwise, determining that the trajectory is normal.
8. A trajectory anomaly route detection system based on self-supervised trajectory representation learning, performing a trajectory anomaly route detection method based on self-supervised trajectory representation learning according to claim 1, characterized in that, The method comprises the following steps: a data acquisition module configured to acquire trajectory data; a preprocessing module configured to preprocess the acquired trajectory data; a learning module configured to construct trajectory dual-view contrastive representation learning based on a dual-view synchronous mask strategy; a fusion module configured to encode time dynamics based on a spatiotemporal fusion mechanism and fuse with spatial features; a reconstruction module configured to learn essential representations from the fused spatiotemporal features for trajectory reconstruction; a judgment module configured to perform error inspection based on the reconstructed trajectory, and determine abnormal trajectories through reconstruction error.
Citation Information
Patent Citations
Double-perspective learning-based mountainous area highway vehicle event detection method
CN105513349A
Map abnormal track detection method and system based on self-attention and unsupervised learning
CN115760918A