A trajectory anomaly detection method based on a bidirectional Mamba model and curriculum learning
Patent Information
- Application Number
- CN202610145627.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-02
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-02-02
AI Technical Summary
(1)基于Transformer(变换器)架构的检测方法在处理长轨迹数据时,因自注意力机制的平方级计算复杂度,导致推理效率低下、硬件资源消耗过大,难以满足实时检测场景需求
本发明通过获取轨迹的具有时间戳的GPS点序列实现对轨迹点信息捕获,其中,每一轨迹采样点包括移动物体在时间戳下的经度和纬度坐标信息,由此可提取得到包括经度、纬度、速度、加速度和运动方向角的空间特征和包括每个轨迹采样点的时间戳和相比于轨迹起始点的时间差的时间特征,从而可实现对所有轨迹点信息的捕获,有效克服了传统基于重构方法因信息利用不充分导致的表征能力欠佳问题,显著增强了轨迹特征的表达能力;另外,本发明为了高效且精准地提取长轨迹的时空特征,适配课程学习中不同难度轨迹样本的特征学习需求,本发明还设计了一种基于双向Mamba模型的轨迹特征编码器,该编码器通过归一化、双线性流拆分、双向一维卷积与状态空间模型的协同设计,既以线性时间复杂度解决了传统方法处理长轨迹的效率瓶颈,又能双向捕捉轨迹的历史与未来上下文信息,结合门控机制筛选关键时空特征,可有效提取常规轨迹的基础模式与异常轨迹的微弱特征;其输出的高质量轨迹特征可直接支撑两阶段重叠课程学习即SC-TOCL策略的分层训练,两者协同可实现长轨迹异常检测的高效性、高精度与高鲁棒性;并且本发明采用双向Mamba编码解码架构,使得模型不仅在计算效率上表现卓越,更凭借其强大的长序列建模能力,精准捕捉了轨迹在双向时空维度下的动态演化规律,实现了从离散观测点到连续移动模式的深度重构;此外,本发明进一步引入空间位置误差和运动状态误差该种基于多步重构误差的投票评分机制,通过多步评估结果提升了异常检测的鲁棒性与判别准确率,为复杂场景下的轨迹分析提供了更具稳定性的技术支撑。再者,本发明创新性地提出了一种基于快照一致性的两阶段重叠课程学习策略,该策略一方面以模型预热快照替代多专家模型,低成本量化轨迹样本难度;另一方面通过简单常规轨迹优先学、复杂异常关联轨迹渐进学、新旧样本重叠复习+全局巩固的逻辑,让模型扎实掌握轨迹正常时空规律,精准捕捉微弱异常特征,同时避免训练遗忘。
Smart Images

Figure CN122196725B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation technology, and in particular to a trajectory anomaly detection method based on a bidirectional Mamba model and curriculum learning. Background Technology
[0002] Vehicle trajectory anomaly detection is one of the core technologies in intelligent transportation systems, autonomous driving safety, and fleet management. Its goal is to automatically identify abnormal trajectories that deviate from normal driving behavior, such as sharp turns, illegal lane changes, and deviations from the path, from massive amounts of vehicle positioning and navigation data. Accurate anomaly detection is crucial for improving road safety, analyzing traffic accidents, and optimizing traffic planning. Although deep learning technology has been widely applied in this field, existing methods, especially when processing long sequences and complex pattern trajectory data, still face several key technical bottlenecks: (1) When processing long trajectory data, the detection method based on the Transformer architecture suffers from low inference efficiency and excessive hardware resource consumption due to the quadratic computational complexity of the self-attention mechanism, making it difficult to meet the requirements of real-time detection scenarios.
[0003] (2) Existing schemes based on the one-way Mamba (a sequence deep learning method) architecture can only capture the one-way temporal dependency of the trajectory, and cannot fully explore the forward and backward correlation features of the trajectory, resulting in insufficient accuracy of anomaly identification and easy to miss or misdetect.
[0004] (3) In the traditional training mode, the model lacks a gradual guidance in learning trajectory data and has limited ability to learn complex trajectory patterns and weak abnormal features, which further restricts the detection performance. Summary of the Invention
[0005] This invention aims to at least partially address the technical problems in related technologies. Therefore, the purpose of this invention is to provide a trajectory anomaly detection method based on a bidirectional Mamba model and curriculum learning. This method improves the completeness of trajectory feature representation through bidirectional feature extraction, ensures efficiency in long trajectory processing by utilizing the linear complexity of the Mamba architecture, optimizes anomaly judgment accuracy through the dual constraints of dynamic time warping and reconstruction error, and guides the model to learn systematically through a curriculum learning strategy. Ultimately, it achieves efficient, high-precision, and highly robust trajectory anomaly detection in long trajectory scenarios.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: A trajectory anomaly detection method based on a bidirectional Mamba model and curriculum learning is implemented through a trajectory anomaly detection model, which includes a spatiotemporal encoder, a bidirectional Mamba encoder, a Mamba decoder, and an anomaly score calculation module connected in sequence. The method includes: A spatiotemporal encoder is used to encode the input trajectory data, extracting the spatial and temporal features of the trajectory and encoding them as spatiotemporal hidden features; The spatiotemporal hidden features are bidirectionally encoded using a Mamba encoder to output the deep hidden state features of the trajectory. The deep hidden state features are reconstructed using a Mamba decoder, and the reconstructed trajectory features are output. An anomaly score calculation module calculates an anomaly score based on the reconstructed trajectory features and the original trajectory features, and determines whether the trajectory is abnormal according to a preset threshold. The original trajectory features are the trajectory data input to the spatiotemporal encoder.
[0007] In one possible implementation, the original trajectory features are obtained through the following steps: A trajectory is obtained, which is a sequence of GPS points with timestamps, wherein each GPS point is a trajectory sampling point, and each trajectory sampling point includes the longitude and latitude coordinates of the moving object under the timestamp; The trajectory features are extracted to obtain the original trajectory features, which include spatial features and temporal features. The spatial features include the longitude, latitude, velocity, acceleration and motion direction angle corresponding to each trajectory sampling point. The temporal features include the timestamp of each trajectory sampling point and the time difference compared to the trajectory starting point.
[0008] In one possible implementation, a spatiotemporal encoder is used to perform feature encoding on the input trajectory data, extracting the spatial and temporal features of the trajectory and encoding them into spatiotemporal hidden features, including: The spatial features of the trajectory are normalized and then mapped to obtain spatial embedding features through a multilayer perceptron. The temporal features of the trajectory are decomposed using a tokenization method to obtain coarse-grained and fine-grained temporal sub-features. A Fourier encoder is used to encode the temporal sub-features, generating coarse-grained and fine-grained temporal embeddings; The spatial embedding feature, coarse-grained temporal embedding, and fine-grained temporal embedding are spliced together to form the spatiotemporal hidden feature.
[0009] In one possible implementation, a bidirectional Mamba encoder is used to bidirectionally Mamba encode the spatiotemporal hidden features, outputting deep hidden state features of the trajectory, including: The spatiotemporal hidden features are normalized. The normalized features are divided into data stream and gated stream by linear layer projection; The data stream is processed by forward and backward one-dimensional convolution to obtain forward convolution features and backward convolution features; The forward and backward convolutional features are processed by forward and backward state space models respectively to obtain the forward SSM output and the backward SSM output. The gated flow is used to generate a gated signal through an activation function; The forward SSM output and the backward SSM output are multiplied element-wise by the gated signal and then added together to obtain the fused feature; The fused features are projected through a linear layer and residually connected with the input spatiotemporal hidden features to obtain the deep hidden state features.
[0010] In one possible implementation, a Mamba decoder is used to reconstruct the deep hidden state features, outputting reconstructed trajectory features, including: The deep hidden state features are input into a decoder consisting of multiple stacked Mamba blocks, each containing normalization, a state-space model, and residual connection operations. The final output of the decoder is mapped back to the original trajectory feature dimension through a linear projection layer to obtain the reconstructed trajectory features, which include reconstructed spatial features and reconstructed temporal features.
[0011] In one possible implementation, an anomaly score calculation module calculates an anomaly score based on the reconstructed trajectory features and the original trajectory features, and determines whether the trajectory is abnormal according to a preset threshold, including: Reconstructed spatial features are extracted from the reconstructed trajectory features, including reconstructed latitude and longitude, velocity, acceleration, and motion direction angle; Extracting original spatial features from original trajectory features; Calculate the squared Euclidean distance between the original latitude and longitude and the reconstructed latitude and longitude of each trajectory sampling point to obtain the spatial position error of each trajectory sampling point; Calculate the velocity error, acceleration error, and motion direction angle error for each trajectory sampling point to obtain the motion state error of each trajectory sampling point; The spatial position error and motion state error are weighted and summed. The outlier score is obtained by averaging the weighted errors of all sampling points. The abnormal score is compared with a preset threshold. If the abnormal score is greater than the preset threshold, the trajectory is determined to be abnormal; otherwise, it is determined to be normal.
[0012] In one possible implementation, the method further includes: pre-training the trajectory anomaly detection model using a two-stage overlapping course learning strategy based on snapshot consistency.
[0013] In one possible implementation, the snapshot-consistent two-stage overlapping course learning strategy specifically includes a snapshot-consistent difficulty assessment stage and a two-stage overlapping course training stage.
[0014] In one possible implementation, the difficulty assessment phase based on snapshot consistency includes: The trajectory anomaly detection model is pre-trained using the full training dataset, the number of pre-training rounds is set, and model snapshots of K model parameters are saved at preset intervals. For each trajectory sample, input to K model snapshots to obtain K representation vectors; Calculate the cosine similarity between each pair of K representation vectors, and take the average as the consistency score. The higher the consistency score, the simpler the trajectory sample. The entire training dataset is sorted from high to low based on consistency scores to obtain an ordered dataset.
[0015] In one possible implementation, the two-stage overlapping course training phase includes: Single-step overlapping progressive training involves dividing an ordered dataset into multiple training rounds, determining the sliding window size and sliding step size, and using a subset of samples within the sliding window to train the trajectory anomaly detection model in each training round, wherein the sample subsets of adjacent rounds have overlapping regions. Global consolidation training, which involves fine-tuning the trajectory anomaly detection model using the full training dataset or a subset of samples uniformly sampled from each difficulty level after progressive training is completed, until the model converges.
[0016] This invention has at least the following technical effects: This invention captures trajectory point information by obtaining a sequence of GPS points with timestamps. Each trajectory sampling point includes the longitude and latitude coordinates of the moving object at the timestamp. This allows for the extraction of spatial features including longitude, latitude, velocity, acceleration, and direction angle, as well as temporal features including the timestamp of each trajectory sampling point and the time difference compared to the trajectory's starting point. This enables the capture of all trajectory point information, effectively overcoming the problem of insufficient representation ability caused by insufficient information utilization in traditional reconstruction methods, and significantly enhancing the expressive power of trajectory features. Furthermore, to efficiently and accurately extract the spatiotemporal features of long trajectories and adapt to the feature learning needs of trajectory samples of varying difficulty in course learning, this invention also designs a trajectory feature encoder based on a bidirectional Mamba model. This encoder, through the collaborative design of normalization, bilinear flow decomposition, bidirectional one-dimensional convolution, and a state-space model, solves the problem of processing long trajectories with linear time complexity, a feature encoding method that is traditionally inefficient. This invention addresses the efficiency bottleneck of trajectory detection while simultaneously capturing historical and future contextual information of the trajectory bidirectionally. Combined with a gating mechanism to filter key spatiotemporal features, it effectively extracts the basic patterns of regular trajectories and the subtle features of anomalous trajectories. The high-quality trajectory features output directly support hierarchical training of the two-stage overlapping course learning, i.e., the SC-TOCL strategy. The synergy of these two approaches achieves high efficiency, high accuracy, and high robustness in long trajectory anomaly detection. Furthermore, the invention employs a bidirectional Mamba encoding and decoding architecture, enabling the model to not only excel in computational efficiency but also, through its powerful long-sequence modeling capabilities, accurately capture the dynamic evolution of the trajectory in the bidirectional spatiotemporal dimension, achieving deep reconstruction from discrete observation points to continuous movement patterns. In addition, this invention further introduces a voting scoring mechanism based on multi-step reconstruction errors, incorporating spatial position error and motion state error. This multi-step evaluation improves the robustness and accuracy of anomaly detection, providing more stable technical support for trajectory analysis in complex scenarios. Furthermore, this invention innovatively proposes a two-stage overlapping course learning strategy based on snapshot consistency. On the one hand, this strategy replaces the multi-expert model with model warm-up snapshots, quantifying the difficulty of trajectory samples at low cost. On the other hand, through the logic of prioritizing simple conventional trajectories, progressively learning complex abnormal correlation trajectories, and overlapping review of new and old samples + global consolidation, the model can solidly grasp the normal spatiotemporal laws of trajectories, accurately capture weak abnormal features, and avoid training forgetting.
[0017] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0018] Figure 1 This is a flowchart of a trajectory anomaly detection method based on a bidirectional Mamba model and course learning, according to an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram of the trajectory anomaly detection model according to an embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of trajectory anomalies according to an embodiment of the present invention.
[0021] Figure 4 This is a schematic diagram of a spatiotemporal encoder according to an embodiment of the present invention.
[0022] Figure 5 This is a schematic diagram of the state space model according to an embodiment of the present invention. Detailed Implementation
[0023] The following describes this embodiment in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.
[0024] The following description, with reference to the accompanying drawings, illustrates a trajectory anomaly detection method based on a bidirectional Mamba model and course learning, according to an embodiment of this invention.
[0025] Figure 1 This is a flowchart of a trajectory anomaly detection method based on a bidirectional Mamba model and course learning, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of a trajectory anomaly detection model according to an embodiment of the present invention. This method is implemented using a trajectory anomaly detection model, such as... Figure 2 As shown, the trajectory anomaly detection model includes a spatiotemporal encoder, a bidirectional Mamba encoder, a Mamba decoder, and an anomaly score calculation module connected in sequence. Figure 1 As shown, the method includes: Step S101: Use a spatiotemporal encoder to perform feature encoding on the input trajectory data, extract the spatial and temporal features of the trajectory, and encode them as spatiotemporal hidden features.
[0026] Step S102: Use a bidirectional Mamba encoder to encode the spatiotemporal hidden features bidirectionally to output the deep hidden state features of the trajectory.
[0027] Step S103: Use the Mamba decoder to reconstruct the deep hidden state features and output the reconstructed trajectory features.
[0028] Step S104: The anomaly score calculation module calculates the anomaly score based on the reconstructed trajectory features and the original trajectory features, and determines whether the trajectory is abnormal according to the preset threshold. The original trajectory features are the trajectory data input to the spatiotemporal encoder.
[0029] This embodiment decouples and deeply fuses the static attributes and dynamic evolution patterns of a trajectory through a spatial encoder and a temporal encoder, forming a unified spatiotemporal hidden feature. This provides a more comprehensive and powerful input foundation for subsequent deep processing. This embodiment also leverages the linear computational complexity of the Mamba architecture's State Space Model (SSM) to overcome the computational bottleneck of Transformer-type models on long sequences, achieving efficient processing of long trajectories. Its bidirectional structure (forward + backward) can simultaneously capture the historical and future context of the trajectory, resulting in a more comprehensive and logically consistent understanding of the behavior of any trajectory point, significantly improving the completeness of feature representation. The core task of the decoder in this embodiment is to reconstruct a normal trajectory from the deep hidden state features. Then, by calculating the difference between the reconstructed trajectory and the original trajectory, an anomaly score is obtained. This eliminates the need for pre-labeling of anomalous samples, enabling accurate detection of trajectory anomalies.
[0030] In one embodiment, the original trajectory features are obtained through the following steps: acquiring a trajectory, for example, Figure 2 The spatiotemporal trajectory T in the model is a sequence of GPS (Global Positioning System) points with timestamps, where each GPS point is a trajectory sampling point. =( , , () indicates the time stamp of the moving object. Longitude coordinates and latitude coordinates ,like Figure 2 In ~ Also note It is a trajectory The original length.
[0031] For anomaly trajectory detection, if the spatiotemporal trajectory T deviates significantly from most expected paths with the same origin and destination, it is considered anomaly. The goal of anomaly trajectory detection is to learn regular path patterns and identify whether each trajectory exhibits abnormal behavior during travel. For example, in Figure 3 In the process, trajectory T3 is considered an anomaly.
[0032] Furthermore, trajectory feature extraction is performed on the trajectory to obtain trajectory extension data, i.e., the original trajectory features. Spatial features , Represent real numbers, The length of the trajectory sequence (i.e., the total number of sampling points in the trajectory) is represented by dimension 5, which corresponds to the longitude (lat), latitude (lon), velocity (speed), acceleration (acc), and direction angle (dir) of each trajectory sampling point. Temporal characteristics. Dimension 2 corresponds to the timestamp of each trajectory sampling point and the time difference compared to the trajectory starting point. Further embedding operations are required for normal input. To accurately encode the temporal and spatial features of the trajectory, the following methods were designed: Figure 4 The time encoder and spatial encoder are shown.
[0033] In one embodiment, a spatiotemporal encoder is used to encode the input trajectory data to extract spatial and temporal features of the trajectory and encode them into spatiotemporal hidden features. This includes: normalizing the spatial features of the trajectory and mapping them through a multilayer perceptron to obtain spatial embedding features; decomposing the temporal features of the trajectory into coarse-grained and fine-grained temporal sub-features; encoding the temporal sub-features using a Fourier encoder to generate coarse-grained and fine-grained temporal embeddings; and concatenating the spatial embedding features, coarse-grained and fine-grained temporal embeddings to form spatiotemporal hidden features.
[0034] In a spatial encoder, normalization is the first step. ,in It is a predefined spatial boundary. Represents the normalized spatial characteristics. This represents the minimum value of a predefined spatial boundary. This represents the maximum value of a predefined spatial boundary. Then it passes through a layer consisting of two linear layers, such as... Figure 4 A multilayer perceptron consisting of linear layer 1 and linear layer 2 is used to extract the core features of the location. The formula used is as follows: (1) in, , These are the weights and biases of the first linear layer, i.e., linear layer 1. , These are the weights and biases of the second linear layer, i.e., linear layer 2. This represents the output dimension of the first linear layer. This indicates the output dimension of the second linear layer. yes Activation function. The spatial dimension of the trajectory extension data is obtained through a spatial encoder. Figure 4 Spatial embedding features .
[0035] In a time encoder, time data is tokenized to acquire time features. It is decomposed into 5 time sub-features: the current time's position within a week, the current time's position within a day, the current time's position within an hour, the minute-level time interval, and the minute-level timestamp, denoted as... Figure 4In Then, a Fourier encoder is used to encode the continuous numerical temporal sub-features into high-dimensional vectors that retain distance / period / scale features. The specific formula is as follows: (2) in, For time sub-features The corresponding high-dimensional vector, and These are the first and second learnable parameters. For embedded dimensions, For the first i Each time feature is further divided into coarse and fine granular dimensions and concatenated, then processed separately. Figure 4 Linear layer projection yields coarse / fine-grained temporal embeddings. The coarse-grained temporal embedding uses... The specific formula is as follows: (3) in, This is a vector concatenation operation. , , They are respectively The corresponding high-dimensional vector. This represents the features after coarse-grained concatenation, and fine-grained temporal embedding is used. The specific formula is as follows: (4) in, , They are respectively The corresponding high-dimensional vector, These are features resulting from fine-grained splicing.
[0036] Then, through linear projection, we obtain... Figure 4 In That is, coarse-grained temporal embedding, and That is, fine-grained temporal embedding (respectively) Figure 2 The coarse-grained and fine-grained temporal features are then combined and concatenated with spatial feature representations to obtain the final trajectory features. Figure 2 Spatiotemporal hidden features : (5) in, , These are the coarse-grained and fine-grained temporal embedding features after linear projection, respectively. Figure 4 In and .
[0037] In one embodiment, a bidirectional Mamba encoder is used to encode the spatiotemporal hidden features bidirectionally to output the deep hidden state features of the trajectory. This includes: normalizing the spatiotemporal hidden features; projecting the normalized features through a linear layer to divide them into a data stream and a gated stream; performing forward and backward one-dimensional convolutions on the data streams to obtain forward convolution features and backward convolution features; processing the forward and backward convolution features using forward and backward state space models respectively to obtain forward SSM output and backward SSM output; generating a gated signal from the gated stream using an activation function; multiplying the forward and backward SSM outputs element-wise with the gated signal and then adding them together to obtain fused features; projecting the fused features through a linear layer and performing a residual connection with the input spatiotemporal hidden features to output the deep hidden state features.
[0038] First, such as Figure 2 As shown, the bidirectional Mamba encoder performs layer normalization on the spatiotemporal hidden features of the input through a normalization layer to stabilize the training distribution.
[0039] (6)
[0040] in, This represents the normalization function. The data after normalization. Projection is performed through two parallel linear layers, expanding the dimensions and dividing them into two main flows: the data flow... x Used for subsequent convolution and SSM processing to extract data patterns, gating flow. z Used for filtering useful information. The specific formula for linear projection is as follows: (7) (8) in, , This represents two linear projection functions. (When entering SSM...) Figure 2 Before the forward and backward state space model in the middle, the data flow x It needs to go through a 1D convolutional layer first, such as Figure 2 The convolution process involves forward and backward convolutions. This step primarily extracts local, neighboring features and provides causal (forward) or anti-causal (backward) inductive biases. The specific formula for the convolution operation is as follows: (9) (10) in, , Data streams x Features after forward and backward 1D convolution It is a smooth nonlinear activation function that can alleviate gradient vanishing and enhance the model's ability to capture complex features. , These are the forward convolution features and backward convolution features after activation, respectively.
[0041] The convolutional features are fed into the forward and backward SSM modules, i.e., the forward and backward state-space models. Discretized state-space equations are utilized here. For each time step in the sequence... t : The input to the forward path is the forward trajectory feature after convolution. First through Figure 5 The input-state mapping module B converts input features into vectors in the state space, injecting the current input information into the state. The previous forward state... First, it goes through state transition module A, which completes the temporal evolution of the historical states according to preset update rules. Then, the historical state evolution results are compared with the input injected state space vector. Figure 5 The addition modules in the code are superimposed to obtain the forward state at the current time. The SSM state update formula is as follows: (11) in, This refers to the state space dimension (the implicit space dimension defined by the model). It is the forward state transition matrix. It is the forward input-state mapping matrix. Figure 5 Modules in f Responsible for the current forward state Transmit the forward state for the next time step. This enables the temporal continuation of states.
[0042] Forward state at the current moment The state-output mapping module C converts the contextual information in the state into output features. Simultaneously, the input features... The input-output direct mapping module D preserves the original details of the input. Finally, these two results are processed through... Figure 5 The addition modules in the middle are merged to obtain the output of the forward path. The corresponding formula is as follows: (12) in, It is the forward state-output mapping matrix. It is a forward input-output direct mapping matrix.
[0043] The backward path typically involves reversing the sequence on the timeline for processing, or using reverse recursive parameters to capture future contextual information. Its specific data processing flow is the same as the forward path, and the specific calculation formula is as follows: (13) (14) in, This is the backward state. This represents the backward state at the next moment. The backward trajectory features are obtained after convolution processing. It is the backward state transition matrix, and its parameters are usually related to... Independent, to adapt to the contextual characteristics of reverse sequences. It is a backward input-state mapping matrix, which maps the time steps in the reverse sequence. t The input features are mapped to the backward state space. It is a backward state-output mapping matrix. It is a backward input-output direct mapping matrix. After the same processing as the forward path, the backward path output is obtained. .
[0044] Meanwhile, the previous gating flow z After an activation function Get gating signal The specific formula is as follows: (15) bidirectional SSM output ( and ) respectively with the gating signal Element-wise multiplication is performed, using a gating mechanism to selectively retain important information. Then, the results of the forward and backward passes are added together. The specific formula is as follows: (16) Features after fusion The output linear layer projects the dimension back to the target dimension. The specific formula is as follows: (17) in, This indicates the output of a linear layer. Finally, the model's output... Spatiotemporal hidden features of the original input Perform residual connection (Add) to obtain the final encoder output. The specific formula is as follows: (18) In one embodiment, a Mamba decoder is used to reconstruct deep hidden state features and output reconstructed trajectory features. This includes: inputting the deep hidden state features into a decoder consisting of multiple stacked Mamba blocks, each Mamba block containing normalization, a state space model, and residual connection operations; and mapping the final output of the decoder back to the original trajectory feature dimensions through a linear projection layer to obtain reconstructed trajectory features, which include reconstructed spatial features and reconstructed temporal features.
[0045] In this embodiment, the main task of the Mamba decoder is to map the highly compressed deep hidden state features output by the encoder back to the original trajectory feature space for reconstructing the input.
[0046] The Mamba decoder receives the final fused feature sequence from the output of the bidirectional Mamba encoder. .in, It is the sequence length. This refers to the hidden layer dimension of the model. A Mamba decoder typically consists of multiple stacked Mamba blocks. Similar to a bidirectional Mamba encoder, each layer contains normalization, an SSM core module, and residual connections, with the aim of gradually recovering local details of the trajectory by utilizing global context information in the hidden states.
[0047] Assuming the Mamba decoder has P layers, for the The layers are: (19) (20) Here, Mamba ( ) represents the Selective State Space Model (SSM) already defined in the bidirectional Mamba encoder, used for feature mixing in the sequence dimension, and FFN represents the fully connected operation. (That is, the output of the bidirectional Mamba encoder is used as the input of the 0th layer of the Mamba decoder). Represents the Mamba decoder's... l Layer input. After processing by the P layer, the final hidden state output of the Mamba decoder is obtained. .
[0048] The output of the last layer of the Mamba decoder The dimension is still the hidden layer dimension. In order to obtain the original trajectory features Reconstructing features with consistent dimensions requires passing through a linear projection layer. The entire decoding process can be summarized as follows: (twenty one) in, Indicates the reconstructed trajectory features, This is a sequence model based on the Mamba state-space model.
[0049] In one embodiment, an anomaly score calculation module calculates an anomaly score based on reconstructed trajectory features and original trajectory features, and determines whether the trajectory is abnormal according to a preset threshold. This includes: extracting reconstructed spatial features from the reconstructed trajectory features, including reconstructed latitude and longitude, velocity, acceleration, and motion direction angle; extracting original spatial features from the original trajectory features; calculating the squared Euclidean distance between the original latitude and longitude and the reconstructed latitude and longitude of each trajectory sampling point to obtain the spatial position error of each trajectory sampling point; calculating the velocity error, acceleration error, and motion direction angle error of each trajectory sampling point to obtain the motion state error of each trajectory sampling point; performing a weighted summation of the spatial position error and motion state error; averaging the weighted errors of all sampling points to obtain the anomaly score; and comparing the anomaly score with a preset threshold. If the anomaly score is greater than the preset threshold, the trajectory is determined to be abnormal; otherwise, it is determined to be normal.
[0050] Traditional trajectory anomaly detection models rely solely on Euclidean distance at a single location to determine trajectory anomalies, easily overlooking latent anomalies with small positional deviations but abnormal velocity / acceleration / direction angles. To address this issue, this embodiment utilizes pure spatial trajectory features reconstructed by a decoder. The system combines latitude and longitude into spatial distance error and incorporates three advanced motion features—velocity, acceleration, and direction angle—to design a weighted comprehensive anomaly score. Normal trajectories must simultaneously satisfy spatial position matching and motion state (velocity, acceleration, and direction) matching. A differentiated weighting method balances the judgment weights of basic position features and advanced motion features. Anomaly score The specific formula for the mean weighted mean square error (MSE) of "spatial distance + high-level motion features" for L trajectory sampling points is as follows: ) (twenty two) in, For the original / reconstructed trajectory t The latitude and longitude vector at any given time is assigned the highest weight as the core of anomaly detection. For the original / reconstructed trajectory t Velocity value at any moment For the original / reconstructed trajectory t acceleration value at any time Original / Reconstructed Trajectory t The angle value of the direction of motion at any moment. , These are the spatial location weights and motion state weights, respectively. =1.
[0051] Based on the validation set trajectory data, the distribution of abnormal scores for normal trajectories is statistically analyzed, and the n0% quantile of the normal trajectory scores is taken as the judgment threshold. (The quantiles can be adjusted according to the accuracy requirements of the scene). If the overall anomaly score of the trajectory to be detected (Score) > If the score is less than or equal to the score, it is considered an abnormal trajectory; if the score is less than or equal to the score, The trajectory is determined to be normal, where n0 is an adjustable percentage parameter.
[0052] In one embodiment, the method further includes pre-training the trajectory anomaly detection model using a two-stage overlapping course learning strategy based on snapshot consistency.
[0053] In this embodiment, the two-stage overlapping course learning strategy based on snapshot consistency specifically includes a difficulty assessment stage based on snapshot consistency and a two-stage overlapping course training stage.
[0054] The difficulty assessment phase based on snapshot consistency includes: pre-training the trajectory anomaly detection model using the full training dataset, setting the number of pre-training rounds, and saving model snapshots of K model parameters at preset intervals; for each trajectory sample, inputting to the K model snapshots to obtain K representation vectors; calculating the cosine similarity between each pair of the K representation vectors, taking the average as the consistency score, with a higher consistency score indicating a simpler trajectory sample; and sorting the full training dataset from high to low according to the consistency score to obtain an ordered dataset.
[0055] The two-stage overlapping course training phase includes: single-step overlapping progressive training and global consolidation training. Single-step overlapping progressive training involves dividing the ordered dataset into multiple training rounds, determining the sliding window size and sliding step size, and using a subset of samples within the sliding window to train the trajectory anomaly detection model in each training round. The sample subsets of adjacent rounds overlap. Global consolidation training involves fine-tuning the trajectory anomaly detection model using the full training dataset or a subset of samples uniformly sampled from each difficulty level after the progressive training is completed, until the model converges.
[0056] Specifically, in order to achieve efficient difficulty assessment and ensure training efficiency while enabling the model to smoothly transition from simple to complex and resist forgetting, this embodiment proposes a two-stage overlapping course learning strategy based on snapshot consistency (SC-TOCL). This method includes two core stages: difficulty assessment based on snapshot consistency and two-stage overlapping course training.
[0057] In the difficulty assessment stage based on snapshot consistency, full-data warm-up is first performed. After initializing the trajectory representation learning model, the model is quickly warm-up trained using the full training dataset, and the number of warm-up epochs is set to Ewarm, so as to ensure that the model initially converges but does not completely overfit. Then model snapshots are saved, and K model snapshots storing model parameters are stored at preset intervals during the warm-up process, which are recorded as a set , M represents a model snapshot set, represents the k -th model snapshot. These model snapshots correspond to the cognitive state of the model on data at different training maturity levels. Next, for each trajectory sample in the full training dataset , input it into K model snapshots to obtain K representation vectors , represents a set of K representation vectors, represents the k -th representation vector. Then the consistency index of these K representation vectors is calculated through average cosine similarity, and the specific formula is: (23) wherein, for the trajectory sample the higher its anomaly score is, the simpler the sample is. Finally, dataset reordering is performed, and the full training dataset is sorted in descending order according to to obtain an ordered dataset , meanwhile, samples with high similarity and low variance are defined as easy samples, and samples with drastic fluctuations and high variance are defined as difficult samples. Wherein, , are respectively the m、n -th representation vector, represents the average cosine similarity calculation function.
[0058] In order to solve the catastrophic forgetting problem and improve the generalization ability of the model through the ordered dataset, a two-stage overlapping curriculum training strategy is introduced, and this processing specifically includes two sub-processes: single-step overlapping progressive training and global consolidation training: (1) In order to enable the model to learn simple patterns first and then gradually introduce complex patterns, and review existing knowledge through data overlapping, the single-step overlapping progressive training process divides into T training epochs, defines window size W and step size S (S < W), and constructs a training subset for the τ-th training epoch , this training subset includes samples starting from index with a length of . and have an overlapping area , This is the training subset constructed for the (τ-1)th training round. This allows the model to review samples that were relatively difficult in the previous round but are relatively easy in the current round when encountering new, more difficult samples. The subset is then used... The model is trained for several rounds. As τ increases, the window slides towards the hardest samples until it covers the entire dataset; (2) In order to avoid forgetting simple basic features and establishing global cognition when the model focuses on difficult samples, the global consolidation training process breaks the course order after completing the single-step overlapping progressive training. The model is fine-tuned until convergence using the full training dataset or a subset uniformly sampled from each difficulty level, so as to ensure that the model integrates the features of simple and difficult samples and forms a robust final representation.
[0059] The trajectory anomaly detection method of this embodiment will be described below as a specific example.
[0060] Data preprocessing and trajectory feature extraction: First, acquire the taxi's GPS trajectory data. The original trajectory is defined as a sequence of GPS points with timestamps. Each point contains longitude, latitude, and a timestamp. .
[0061] Feature extraction is performed on the original trajectory to generate a preliminary trajectory feature representation, namely the original trajectory features. Spatial features Time characteristics These spatial features can include physical characteristics such as velocity, acceleration, or azimuth angle, which can serve as inputs for subsequent models.
[0062] Spatio-Temporal Encoder: Spatial features of the trajectory The input spatial encoder first performs normalization, and then maps the input spatial embedding features through a multilayer perceptron (MLP). .
[0063] Time characteristics of the trajectory The input time encoder first performs tokenization processing, decomposing the timestamp into coarse-grained time sub-features. and fine-grained temporal sub-features Then, a Fourier encoder is used to utilize the learnable parameters. and The specific formula for generating the time embedding by performing a cosine transform is as follows: (twenty four) Subsequently, coarse-grained temporal embeddings were obtained through linear layer projection. and fine-grained temporal embedding .
[0064] Finally, the spatial embedding features, fine-grained temporal embeddings, and coarse-grained temporal embeddings are concatenated to generate a unified spatiotemporal latent feature for the trajectory. .
[0065] Bi-Mamba Encoder: Hidden features of spacetime Input a three-layer stacked bidirectional Mamba encoder to capture long sequence dependencies with linear complexity. The first input... After layer normalization (RMSNorm), the data stream is divided into data streams through linear layer projection. x and gating .
[0066] Data Stream The inputs then proceed to two branches: Forward and Backward. Each branch first extracts local neighborhood features through a 1D convolutional layer (Conv1D). The features after convolution... and The forward SSM and backward SSM modules are entered respectively. The discretized state-space equations, namely formulas (11) and (12), are used to capture the forward and backward features of the trajectory, respectively, to obtain the forward SSM output. and backward SSM output .
[0067] Gated flow The gating signal is obtained after the SiLU activation function. . and Each element is multiplied element-wise by the gate signal and then added together to obtain the result. .
[0068] Fusion features After projection onto the output linear layer and residual connection with the input, the deep hidden state features are output. .
[0069] Trajectory Reconstruction and Decoding (Mamba Decoder):
[0070] The Mamba decoder is constructed from three stacked Mamba blocks. The decoder receives deep hidden state features from the encoder's output. The trajectory details are gradually recovered using the global context capability of the SSM (Straight Path Model). Finally, a linear projection layer maps the dimensions back to the original feature space, resulting in the reconstructed trajectory feature sequence. Then, the features are further broken down to form reconstructed spatial features. and reconstruction time characteristics .
[0071] Training Process Based on Course Learning: To address the challenges of difficult model training and easy forgetting, this embodiment employs a training method based on snapshot consistency and two-stage overlapping course learning: In the difficulty assessment phase based on snapshot consistency, the model is first warmed up with 1000 trajectories using 5 rounds of full-data pre-training to ensure initial convergence. Subsequently, model snapshots of the model parameters from 3 different training rounds are stored. Next, each trajectory sample in the dataset Input three model snapshots and get three representation vectors. And the mean cosine similarity is calculated using formula (23) as the difficulty score, and finally based on Sort the entire training dataset from highest to lowest quality to obtain an ordered dataset. .
[0072] In the two-stage overlapping course training phase, single-step overlapping progressive training is first performed: the ordered dataset is... The training is divided into 99 rounds, with a window size W=20 and a step size S=10. The first round of training is a subset of... Twenty simple trajectories with indices 0-19 were selected as the training subset for the second round. Select samples from index 10-29 (containing 10 records with...) (Overlapping trajectories and 10 new trajectories), and in each subsequent round, the window slides 10 indices toward the difficult samples until all 1000 trajectories are covered. Each subset is trained for 3 rounds to ensure that the model can still review existing knowledge when encountering new and more difficult samples. After progressive training is completed, global consolidation training is performed, and the model is fine-tuned for 2 rounds using all 1000 trajectories. The final output model with weights saved as a .pth file avoids the model forgetting simple basic features and establishes global cognition.
[0073] Abnormal score calculation: Anomalies in a trajectory are determined by the error between the original trajectory and the reconstructed trajectory. If a normal trajectory has a minimal error between the original and reconstructed trajectories after model encoding and decoding, it can be assumed that the model has learned the movement pattern of that path and can fill in the corresponding trajectory; thus, the trajectory is considered normal. Conversely, if an abnormal trajectory exists, and the difference between the reconstructed trajectory and the abnormal trajectory after model encoding and decoding is significant, it can be assumed that the model has failed to learn the corresponding pattern from historical data. Therefore, this trajectory is considered abnormal.
[0074] Abnormal threshold Based on validation set calibration, the 95th percentile of the normal trajectory score is taken. In the example... , =0.75, The abnormal score (Score) is calculated using formula (22). If the score is greater than 0.75, it is determined to be an abnormal trajectory (such as speeding and driving against traffic on a regular route, rapid acceleration and deceleration, or detour trajectory with significant deviation in latitude and longitude). If the score is less than or equal to 0.75, it is determined to be a normal trajectory (the position and motion state are both consistent with the regular pattern).
[0075] In summary, this invention captures trajectory point information by obtaining a sequence of GPS points with timestamps. Each trajectory sampling point includes the longitude and latitude coordinates of the moving object at the timestamp. This allows for the extraction of spatial features including longitude, latitude, velocity, acceleration, and direction angle, as well as temporal features including the timestamp of each trajectory sampling point and the time difference compared to the trajectory's starting point. This enables the capture of all trajectory point information, effectively overcoming the problem of insufficient representation ability caused by inadequate information utilization in traditional reconstruction methods, and significantly enhancing the expressive power of trajectory features. Furthermore, to efficiently and accurately extract the spatiotemporal features of long trajectories and adapt to the feature learning needs of trajectory samples of varying difficulty in course learning, this invention also designs a trajectory feature encoder based on a bidirectional Mamba model. This encoder, through the collaborative design of normalization, bilinear flow decomposition, bidirectional one-dimensional convolution, and a state-space model, solves the problem of traditional methods in linear time complexity. This invention addresses the efficiency bottleneck of long trajectories while simultaneously capturing historical and future contextual information of the trajectory bidirectionally. Combined with a gating mechanism to filter key spatiotemporal features, it effectively extracts the basic patterns of regular trajectories and the subtle features of anomalous trajectories. The high-quality trajectory features output directly support hierarchical training of the two-stage overlapping course learning strategy (SC-TOCL). This synergy achieves high efficiency, high accuracy, and high robustness in long trajectory anomaly detection. Furthermore, the invention employs a bidirectional Mamba encoding and decoding architecture, resulting in superior computational efficiency and, thanks to its powerful long-sequence modeling capabilities, accurately capturing the dynamic evolution of the trajectory in the bidirectional spatiotemporal dimension, achieving deep reconstruction from discrete observation points to continuous movement patterns. In addition, this invention introduces a voting scoring mechanism based on multi-step reconstruction errors, incorporating spatial position error and motion state error. This multi-step evaluation improves the robustness and accuracy of anomaly detection, providing more stable technical support for trajectory analysis in complex scenarios. Furthermore, this invention innovatively proposes a two-stage overlapping course learning strategy based on snapshot consistency. On the one hand, this strategy replaces the multi-expert model with model warm-up snapshots, quantifying the difficulty of trajectory samples at low cost. On the other hand, through the logic of prioritizing simple conventional trajectories, progressively learning complex abnormal correlation trajectories, and overlapping review of new and old samples + global consolidation, the model can solidly grasp the normal spatiotemporal laws of trajectories, accurately capture weak abnormal features, and avoid training forgetting.
[0076] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A trajectory anomaly detection method based on a bidirectional Mamba model and course learning, characterized in that, The method is implemented through a trajectory anomaly detection model, which includes a spatiotemporal encoder, a bidirectional Mamba encoder, a Mamba decoder, and an anomaly score calculation module connected in sequence. The method includes: A spatiotemporal encoder is used to encode the input trajectory data, extracting the spatial and temporal features of the trajectory and encoding them as spatiotemporal hidden features; The spatiotemporal hidden features are bidirectionally encoded using a Mamba encoder to output the deep hidden state features of the trajectory. The deep hidden state features are reconstructed using a Mamba decoder, and the reconstructed trajectory features are output. An anomaly score calculation module calculates an anomaly score based on the reconstructed trajectory features and the original trajectory features, and determines whether the trajectory is abnormal according to a preset threshold. The original trajectory features are the trajectory data input to the spatiotemporal encoder; the trajectory is a sequence of GPS points with timestamps, where each GPS point is a trajectory sampling point, and each trajectory sampling point includes the longitude and latitude coordinates of the moving object under the timestamp. The method further includes: pre-training the trajectory anomaly detection model using a two-stage overlapping course learning strategy based on snapshot consistency; The two-stage overlapping course learning strategy based on snapshot consistency specifically includes a difficulty assessment stage based on snapshot consistency and a two-stage overlapping course training stage. The difficulty assessment phase based on snapshot consistency includes: The trajectory anomaly detection model is pre-trained using the full training dataset, the number of pre-training rounds is set, and model snapshots of K model parameters are saved at preset intervals. For each trajectory sample, input to K model snapshots to obtain K representation vectors; Calculate the cosine similarity between each pair of K representation vectors, and take the average as the consistency score. The higher the consistency score, the simpler the trajectory sample. The entire training dataset is sorted from high to low based on the consistency score to obtain an ordered dataset. The two-stage overlapping course training phase includes: Single-step overlapping progressive training involves dividing an ordered dataset into multiple training rounds, determining the sliding window size and sliding step size, and using a subset of samples within the sliding window to train the trajectory anomaly detection model in each training round, wherein the sample subsets of adjacent rounds have overlapping regions. Global consolidation training, which involves fine-tuning the trajectory anomaly detection model using the full training dataset or a subset of samples uniformly sampled from each difficulty level after progressive training is completed, until the model converges.
2. The method as described in claim 1, characterized in that, The original trajectory features are obtained through the following steps: The trajectory features are extracted to obtain the original trajectory features, which include spatial features and temporal features. The spatial features include the longitude, latitude, velocity, acceleration and motion direction angle corresponding to each trajectory sampling point. The temporal features include the timestamp of each trajectory sampling point and the time difference compared to the trajectory starting point.
3. The method as described in claim 1, characterized in that, A spatiotemporal encoder is used to perform feature encoding on the input trajectory data, extracting the spatial and temporal features of the trajectory and encoding them into spatiotemporal hidden features, including: The spatial features of the trajectory are normalized and then mapped to obtain spatial embedding features through a multilayer perceptron. The temporal features of the trajectory are decomposed using a tokenization method to obtain coarse-grained and fine-grained temporal sub-features. A Fourier encoder is used to encode the temporal sub-features, generating coarse-grained and fine-grained temporal embeddings; The spatial embedding feature, coarse-grained temporal embedding, and fine-grained temporal embedding are spliced together to form the spatiotemporal hidden feature.
4. The method as described in claim 1, characterized in that, The spatiotemporal hidden features are bidirectionally encoded using a Mamba encoder to output the deep hidden state features of the trajectory, including: The spatiotemporal hidden features are normalized. The normalized features are divided into data stream and gated stream by linear layer projection; The data stream is processed by forward and backward one-dimensional convolution to obtain forward convolution features and backward convolution features; The forward and backward convolutional features are processed by forward and backward state space models respectively to obtain the forward SSM output and the backward SSM output. The gated flow is used to generate a gated signal through an activation function; The forward SSM output and the backward SSM output are multiplied element-wise by the gated signal and then added together to obtain the fused feature; The fused features are projected through a linear layer and residually connected with the input spatiotemporal hidden features to obtain the deep hidden state features.
5. The method as described in claim 1, characterized in that, The deep hidden state features are reconstructed using a Mamba decoder, and the reconstructed trajectory features are output, including: The deep hidden state features are input into a decoder consisting of multiple stacked Mamba blocks, each containing normalization, a state-space model, and residual connection operations. The final output of the decoder is mapped back to the original trajectory feature dimension through a linear projection layer to obtain the reconstructed trajectory features, which include reconstructed spatial features and reconstructed temporal features.
6. The method as described in claim 2, characterized in that, An anomaly score calculation module calculates an anomaly score based on the reconstructed trajectory features and the original trajectory features, and determines whether the trajectory is abnormal according to a preset threshold, including: Reconstructed spatial features are extracted from the reconstructed trajectory features, including reconstructed latitude and longitude, velocity, acceleration, and motion direction angle; Extracting original spatial features from original trajectory features; Calculate the squared Euclidean distance between the original latitude and longitude and the reconstructed latitude and longitude of each trajectory sampling point to obtain the spatial position error of each trajectory sampling point; Calculate the velocity error, acceleration error, and motion direction angle error for each trajectory sampling point to obtain the motion state error of each trajectory sampling point; The spatial position error and motion state error are weighted and summed. The outlier score is obtained by averaging the weighted errors of all sampling points. The abnormal score is compared with a preset threshold. If the abnormal score is greater than the preset threshold, the trajectory is determined to be abnormal; otherwise, it is determined to be normal.
Citation Information
Patent Citations
Semantic fingerprint adaptive training method for teaching service robot
CN120653994A
Trajectory abnormal route detection method and system based on self-supervised trajectory representation learning
CN120724351A