Traffic track anomaly detection method and device based on neighborhood reconstruction and graph contrast learning

By abstracting the traffic trajectory data into a graph structure and using graph neural network and graph comparison learning technology, the limitations of traditional methods when processing complex traffic data are solved, efficient traffic trajectory anomaly detection is achieved, and the accuracy and adaptability of detection are improved.

CN119939465AActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510013887.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Traditional traffic trajectory anomaly detection methods face the challenges of high complexity and dynamic changes when processing large-scale traffic data, and it is difficult to effectively identify diversified anomaly patterns and real-time changing traffic flows.

Method used

Using a method based on neighborhood reconstruction and graph comparison learning, the traffic trajectory data is abstracted into graph structure data, the vehicle trajectory diagram structure is reconstructed through graph neural network, and the difference between normal trajectories and abnormal trajectories in the traffic trajectory diagram data is captured through graph comparison learning network to achieve abnormal recognition.

Benefits of technology

It improves the accuracy and real-time detection of traffic trajectory anomaly, can flexibly adapt to changes in different traffic scenarios, identify complex abnormal behaviors and patterns, and enhances the safety and fluency of the traffic system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939465A_ABST
    Figure CN119939465A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic trajectory anomaly detection method and device based on neighborhood reconstruction and graph contrast learning, and the method comprises the steps: constructing a traffic trajectory network through selecting a Porto data set and carrying out data preprocessing, abstracting the trajectory data of each vehicle into nodes in a graph, and representing the relation between vehicles through edges; through a neighborhood reconstruction module, a graph neural network is utilized to encode a receiving domain of a node, a neighborhood structure of the node is reconstructed in a dimension reduction space, and a complex relation and attribute information between vehicles in a traffic track are captured; the positive sample similarity is maximized and the negative sample similarity is minimized by utilizing graph contrast learning so as to enhance the learning ability of the model for the traffic track abnormal behavior, and thus the detection accuracy of the abnormal track is improved; calculating an abnormal score of each track through a defined abnormal scoring function, and setting a threshold value of the abnormal score based on means such as historical data and error analysis; and finally, through ranking abnormal scores, the tracks with high scores are marked as possible abnormities, and effective detection of traffic track abnormal behaviors is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data mining and artificial intelligence, and in particular to a method for detecting anomalies in traffic trajectories, in particular to a method and device for detecting anomalies in traffic trajectories based on neighborhood reconstruction and graph contrast learning. Background Art

[0002] In the context of the current intelligent transportation system and digital transformation, traffic trajectory anomaly detection has become an important tool for traffic management and public safety maintenance. With the continuous increase in urban traffic flow and the surge in the number of vehicles, the transportation system is facing complex and hidden abnormal behaviors such as traffic accidents, traffic congestion, and traffic violations. These problems not only threaten public safety, but may also lead to reduced social efficiency, waste of resources, and even legal liability. If the traffic management department fails to detect and respond to these abnormal behaviors in a timely manner, it may lead to more serious accidents or safety hazards. Therefore, the importance of traffic trajectory anomaly detection is becoming more and more prominent. Traditional detection methods face the challenges of high complexity and dynamic changes when processing large-scale traffic data. The lack of a powerful and intelligent anomaly detection system may cause some malicious behaviors or traffic violations to be missed, further endangering the safety and smoothness of the transportation system. Therefore, it is imperative to build a powerful and intelligent traffic trajectory anomaly detection mechanism. By effectively detecting and preventing the occurrence of traffic accidents, traffic violations, and abnormal traffic patterns, the safety of the transportation system can be effectively maintained, the efficiency of traffic management can be improved, the safety of citizens' travel can be guaranteed, and the social problems caused by traffic accidents can be reduced. This not only helps to ensure the healthy operation of the transportation system, but also creates a safer and more reliable transportation environment for society.

[0003] With the continuous growth of urban traffic data and the increasing complexity of data, traffic trajectory anomaly detection faces the following major challenges:

[0004] 1) High dimensionality and sparsity of data: Traffic trajectory data usually has high-dimensional characteristics, such as vehicle speed, location, driving route, etc. At the same time, the vehicle's driving trajectory is relatively sparse in certain time periods and regions, which makes it difficult for traditional feature statistics-based methods to fully characterize data characteristics.

[0005] 2) Diverse abnormal patterns: Abnormal behaviors in traffic trajectories may manifest themselves in various patterns such as illegal lane changes, speeding, and driving in the wrong direction. They are diverse and hidden, which increases the difficulty of detection.

[0006] 3) Dynamics of traffic flow and trajectory: Traffic trajectories are highly dynamic, and the routes, traffic flow, and speed of vehicles change at any time. Traditional static detection methods are difficult to adapt to these real-time changes, so more dynamically adaptive algorithms are needed.

[0007] 4) Computing resource limitations: Processing large-scale traffic trajectory data requires efficient algorithm design and computing resources, otherwise the practical application may be affected due to excessively high computing costs.

[0008] In general, traditional traffic trajectory anomaly detection methods have certain limitations when facing complex and changing traffic environments, and more adaptive and intelligent methods are urgently needed to improve the accuracy and real-time performance of anomaly detection. New traffic trajectory anomaly detection methods should be able to flexibly adapt to changes in different traffic scenarios, process high-dimensional and heterogeneous data through intelligent algorithms, and more comprehensively and accurately identify bad or abnormal behaviors involving multiple vehicles or traffic points, thereby better maintaining the safety and smoothness of the traffic system.

[0009] Thanks to the development of graph neural network (GNN) technology, traffic trajectory anomaly detection has ushered in new breakthroughs. GNN can effectively capture the complex relationships between different vehicles and traffic points (such as intersections and road sections) in the traffic system, and more comprehensively analyze the dynamic changes of traffic flow, paths and behaviors. Its powerful feature learning ability enables the detection system to adapt to the ever-changing traffic anomaly patterns, which is more flexible and intelligent than traditional methods. By learning the spatiotemporal associations and driving trajectories between vehicles, GNN can identify abnormal behaviors in traffic trajectories, such as speeding, illegal lane changes, and driving in the wrong direction. In addition, GNN helps detect potential traffic anomalies and promptly discover traffic safety hazards and emergencies through the embedded representation of nodes and edges. These technological innovations have promoted the advancement of traffic trajectory anomaly detection and provided a powerful tool for intelligent traffic management. By improving the safety of the transportation system and reducing the impact of traffic accidents and violations on society, it provides a safer and smoother traffic environment for cities. Summary of the invention

[0010] Aiming at the complex data of traffic trajectories, the present invention provides a traffic trajectory anomaly detection method and device based on neighborhood reconstruction and graph contrast learning to overcome the above-mentioned shortcomings of the prior art, so as to realize the accurate detection of abnormal data in the traffic trajectory field.

[0011] The present invention first abstracts the traffic trajectory data into graph structure data, in which nodes represent vehicles or traffic points (such as intersections, road sections), and edges represent the relationship between vehicles or traffic flow. The present invention introduces a neighborhood reconstruction module, which uses the feature information of vehicles and their surrounding traffic points to reconstruct the traffic trajectory graph structure, and evaluates potential abnormal vehicles or traffic points through reconstruction errors. Secondly, positive and negative sample pairs are generated by subgraph sampling so that the model can learn the patterns of normal and abnormal behaviors. Then, the subgraph data is input into the graph convolutional network (GCN) layer to obtain the hidden layer information and potential representation of the subgraph, and gradually extract the features of nodes and edges. Then, the three comparison methods of graph comparison learning (node ​​and node comparison, node and subgraph comparison, subgraph and subgraph comparison) are used to learn the differences between normal and abnormal trajectories in the traffic trajectory graph data. Finally, the abnormal information learned from each comparison mode is integrated for abnormal identification. During the training process, the model judges positive and negative samples by comparing the similarity of the embedded vectors, and is optimized by the triple loss function, etc. to ensure that the similarity of the positive sample pairs is higher than that of the negative sample pairs. Ultimately, the system can identify abnormal behaviors (such as speeding, illegal lane changes, driving in the wrong direction, etc.) or abnormal traffic patterns in the huge and complex traffic trajectory data.

[0012] The present invention achieves the above-mentioned purpose through the following technical solutions: A traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning includes the following implementation steps:

[0013] S1: Select a widely used traffic trajectory dataset, preprocess the raw data, and define the corresponding graph structure data;

[0014] S2: Reconstructing the neighborhood information of graph structured data to capture the similarities and differences in attribute space;

[0015] S3: Subgraph and node sampling to obtain positive and negative sample pairs and extract the potential representation of the subgraph embedding vector;

[0016] S4: Graph contrastive learning networks capture similarities and differences in embedding vectors;

[0017] S5: define anomaly scoring function;

[0018] S6: Perform traffic trajectory data anomaly detection.

[0019] Wherein, step S1 specifically includes:

[0020] S1.1: Selecting a suitable traffic trajectory dataset is crucial for anomaly detection. In order to simulate real-world traffic behavior, the Porto dataset was selected, which is derived from taxi trajectory data in Porto, Portugal. In the dataset, vehicles are regarded as nodes, and the vehicle's location information (latitude and longitude), speed, timestamp, etc. are used as node features. The movement trajectory and intersection relationship between vehicles are represented as edges in the graph. This dataset contains a large amount of taxi driving trajectory data, covering traffic flows in different time periods, regions, and road types. With these data, regular and abnormal driving patterns in the traffic system can be effectively simulated, which helps to identify potential traffic violations or abnormal patterns, such as speeding, driving in the wrong direction, and illegal lane changes. This dataset provides rich actual traffic behavior information for traffic trajectory anomaly detection and is an important reference for traffic behavior anomaly detection.

[0021] S1.2: Data cleaning. When processing the Porto dataset, first remove noise points and redundant points. These points may be caused by weak signals of the positioning system or GPS accuracy problems. They not only increase the scale of data storage, but may also have a significant impact on the accuracy of anomaly detection. Next, handle missing values ​​and fill in missing node feature data. You can use mean interpolation or fill in by analyzing the characteristics of similar vehicles to ensure that the trajectory data of each vehicle is complete and consistent. Finally, perform data conversion to convert non-numerical features into numerical features for easy model processing. This includes feature encoding operations, such as converting timestamps into time intervals, or converting classification features such as different road types and traffic conditions into numerical features to facilitate subsequent model training and analysis.

[0022] S1.3: Construct graph-structured data. When processing the Porto dataset, the traffic trajectory data is abstracted into a graph structure. The trajectory data of each taxi is regarded as a node in the graph, and the characteristics of the node include the vehicle's location information (latitude and longitude), speed, timestamp, etc. The edges in the graph represent the relationship between vehicles, and are usually composed of the movement trajectories of vehicles traveling on the same road section or between adjacent intersections in the same time period. By using road networks (such as intersections and road sections) as nodes and the relationships between vehicles as edges, graph-structured data based on traffic flow can be established. In addition, the weight of the edge can be defined according to the distance or travel time between vehicles, so as to better reflect the dynamic characteristics of the traffic network. Ultimately, this graph-structured data provides an effective input for subsequent anomaly detection and pattern recognition.

[0023] S1.4: Define graph data and anomaly detection problem. For a given undirected graph G = (V, E), where {V 1 , V 2 , V 3 ……V N} represents the set of nodes, N represents the number of nodes, and E represents the set of edges. In addition, the node feature matrix X∈R N×D Represents node feature information, adjacency matrix A∈R N×N Indicates the graph structure information. At the same time, use x i ∈R D Represents node v i The characteristics of d i Represents the degree of each node. For the adjacency matrix A, if A ij =1, it means node v i and v j There is an edge between them, otherwise A ij =0. The goal of the present invention is to detect all abnormal nodes in a given graph. The solution to this problem is as follows: G = (V, E), whose adjacency matrix is ​​A, and the model measures the degree of abnormality of each node in G by learning the abnormality scoring function S(·). S(v i ) value is larger, the node v i The higher the probability of abnormality. Then sort the abnormality scores of all nodes in descending order, and determine the abnormality by selecting a certain threshold ρ;

[0024] S1.5: Divide the processed data into training samples and test samples. This division helps to use part of the data to learn the normal traffic trajectory network data pattern in the model training phase, and use independent data to evaluate the performance of the model in the test phase to ensure the generalization ability and reliability of the model.

[0025] Wherein, step S2 specifically includes:

[0026] S2.1: Reconstructing neighborhood information of graph-structured data. In traffic trajectory anomaly detection, the neighborhood reconstruction module uses a graph neural network (GNN) to encode the receptive field of the vehicle trajectory network and reconstruct the neighborhood structure of the vehicle in the reduced-dimensional space. This reconstruction process not only restores the characteristic attributes of the vehicle itself (such as position, speed, acceleration, etc.), but also reconstructs the connection pattern between vehicles and their traffic relationship with adjacent vehicles, thereby effectively capturing the abnormal information of traffic trajectory data in the attribute space. This method not only uses the representation of vehicle nodes to reconstruct local neighborhood information, but also introduces the structural information of the global traffic network, and further enhances the modeling ability of complex relationships and attribute spaces between vehicles through comparative learning. In this way, the abnormal driving patterns of vehicles in the graph, such as speeding, sudden braking or reverse driving, can be more comprehensively captured, improving the performance of the model in traffic trajectory anomaly detection.

[0027] S2.2: Capturing similarities and differences in attribute space. First, the node’s own representation is updated by iteratively aggregating the node’s neighbor information through the GNN Encoder. This aggregation operation can capture the structural information of the node’s local neighborhood and the relationship characteristics with its direct neighbors. The mathematical expression of this process is as follows:

[0028]

[0029] Among them, f i (l) represents the feature vector of node i in layer l, UPDATE represents the operation used to update node features, and the Aggregation function is used to aggregate the information of neighboring nodes. i Represents the set of all neighbor nodes of node i.

[0030] Through a multi-layer perceptron (MLP) i (l+1) Decoded step by step In this way, the original feature vector f is reconstructed i (0) . Then, use l2-loss to calculate With f i (0) The node reconstruction loss is obtained by the difference between:

[0031]

[0032] Reconstructing node degrees using an MLP Its loss function is as follows:

[0033]

[0034] Neighborhood Experience distribution Using a multivariate Gaussian distribution To approximate, the mean estimation and covariance matrix estimation formulas are as follows:

[0035]

[0036] Then, from f i (l+1) Constructing an approximate distribution Specifically, using f i (l+1) Generate the mean and covariance matrix of the multivariate Gaussian distribution using the following formula:

[0037]

[0038] From the generated Gaussian distribution k samples are sampled from the dataset, and then these samples are transformed into approximate samples represented by neighbor features through a fully connected neural network (FNN) Finally, use the generated neighbor feature sample To estimate the new mean and covariance matrix:

[0039]

[0040] Based on the given and The KL divergence between these two distributions is used to measure the reconstruction loss of neighbor attribute features:

[0041]

[0042] Where p represents the dimension.

[0043] Finally, the total loss in the neighborhood reconstruction module is as follows:

[0044]

[0045] Wherein, step S3 specifically includes:

[0046] S3.1: Extracting the embedding vector of the subgraph is a crucial step in traffic trajectory anomaly detection, which mainly adopts the method of graph convolutional network (GCN) layer. Before this, in order to make the obtained representation more discriminative, it is necessary to mask the features of the target nodes in the subgraph. Their hidden layer feature representation can be expressed by the following formula:

[0047]

[0048] in represents the symmetric normalized adjacency matrix, represents the hidden representation of the lth layer, W (l) Represents weight.

[0049] The GCN layer performs convolution operations on the subgraph, aggregates node features and updates its representation. After multiple layers of GCN are stacked, the node features are gradually improved and the structural information of the subgraph is enriched.

[0050] S3.2: The potential representation of the entire subgraph is obtained by pooling or aggregating the node embedding vectors output by the GCN layer. This process integrates the structural and feature information of the subgraph into the embedding vector, providing a higher-level and more comprehensive subgraph representation for subsequent graph contrast learning, enabling the model to more accurately capture patterns and abnormal behaviors in traffic trajectory data. Then, the final representation of the subgraph is calculated using skip connections, which can increase the connectivity of the nodes in the graph. (l) and a projection vector α (l)Sort the nodes and find the indices corresponding to the first k largest values:

[0051]

[0052] Then, in the original adjacency matrix Extract the corresponding subgraph adjacency matrix from

[0053]

[0054] Furthermore, element-wise matrix multiplication is performed to obtain the new feature matrix:

[0055]

[0056] Finally, the original structure of the graph is restored through the distribute(·) operation and the final representation of the subgraph z is obtained i , the formula is as follows:

[0057] z i =distribute(0 n×c ,X (l+1) ,idx) (13)

[0058] Accordingly, MLP is used to transform the target node features into the same embedding space as the subgraph, and the final representation of the node is obtained. i , and shares weight W with the previous GCN (l) :

[0059] e i =σ(X (l) W (l) ) (14)

[0060] Wherein, step S4 specifically includes:

[0061] S4.1: Based on the above operations, the embedding representation of the positive and negative sample pairs is obtained, and the similarities and differences between them need to be further measured. The process of capturing the similarity and difference of the embedding vectors by the graph contrast learning network is mainly achieved by comparing the similarity of the positive and negative sample pairs. For each pair of positive and negative samples, by calculating the similarity of their embedding vectors, the network can learn that the similarity of the positive sample pairs is higher, while the similarity of the negative sample pairs is lower. A bilinear model is used to measure the relationship between them, which is calculated by the following formula:

[0062]

[0063] Among the positive sample pairs, the target node and the subgraph tend to be similar, that is, s i = 1, on the contrary, in the negative sample pair, the target node and the subgraph may not be similar, that is, s i= 0. Therefore, the contrast loss is calculated using the binary cross entropy loss function:

[0064]

[0065] At the same time, it is also necessary to compare between nodes, which is more conducive to discovering node-level anomalies. Similarly, the features of the target node are shielded, positive and negative sample pairs are constructed for the target node, and the subgraph representation is obtained using the new GCN layer:

[0066]

[0067] Where W ′(l) Different from the parameter matrix of the node-subgraph.

[0068] The node features are mapped to the same hidden space through MLP to obtain the final representation of the node Then a bilinear model is used to evaluate the relationship between nodes. Then the node-node contrast loss function can be defined as:

[0069]

[0070] In order to optimize the embedded representation of subgraphs and learn their intrinsic characteristics, a third comparison mode is defined: subgraph-subgraph comparison. The subgraph in another view that has the same target node As a positive sample pair, the subgraph generated in the two views where the other node is located and is considered as a negative sample pair. The following loss function is used to optimize the contrast:

[0071]

[0072] Finally, the loss in the contrastive learning module is as follows:

[0073]

[0074] Where β∈(0,1) is a parameter used to balance node-level abnormal information, is a parameter used to balance the abnormal information at the subgraph level.

[0075] The GCN layer performs convolution operations on the subgraph, aggregates node features and updates its representation, thereby capturing the relationship between nodes in the traffic trajectory network. By stacking multiple layers of GCN, node features are gradually enhanced to help identify abnormal patterns.

[0076] Wherein, step S5 specifically includes:

[0077] S5.1: By minimizing the set objective function, the anomaly score of each node can be calculated. By minimizing the objective function, the anomaly score of each node can be calculated. In the neighborhood reconstruction module, the anomaly score is calculated based on the node self-reconstruction error. The reconstruction error is measured by the difference between the reconstructed feature vector and the original feature vector. The calculation method is:

[0078]

[0079] Correspondingly, since normal nodes are usually similar to positive samples and dissimilar to negative samples, abnormal nodes are very different from both positive and negative samples. in and Represent the similarity of positive and negative pairs respectively. Based on this rule, the anomaly score of the comparison module can be obtained:

[0080]

[0081] In summary, the two anomaly scores are combined to obtain the anomaly score of each node:

[0082]

[0083] In order to mitigate the inherent randomness of a single detection, multiple detections are performed on each node to obtain their anomaly scores, and the average is taken as the final anomaly score. The above formula is further rewritten as:

[0084]

[0085] Where R represents the number of anomaly detections.

[0086] Wherein, step S6 specifically includes:

[0087] S6.1: In the previous step, the anomaly score of each node was obtained Next, setting a suitable anomaly score threshold is a key step, which directly determines whether the node is judged as abnormal. The selection of the threshold should take into account business needs, system performance and risk tolerance, and is usually determined by the performance indicators of the verification sample or test sample (such as ROC-AUC). Common threshold determination methods include empirical rules, historical data analysis and dynamic adjustment. In the present invention, a certain anomaly score is used as the boundary. If it exceeds the boundary, it is abnormal, otherwise it is normal, that is, the threshold Threshold is obtained:

[0088] Threshold=POT(Score Train ,Score Test ,δ) (25)

[0089] In actual operation, the POT (extreme value theory) method is used to dynamically adjust the threshold according to the estimated risk value. By calculating the anomaly score of each vehicle and forming a sorted list, vehicles with higher anomaly scores are marked as potential abnormal vehicles, helping traffic management personnel focus on high-risk behaviors and improving the efficiency of identifying and processing abnormal driving patterns. This method combines the graph structure information of the traffic network with deep learning technology to effectively enhance the ability to detect traffic trajectory anomalies, and can promptly detect potential traffic violations or abnormal driving patterns, such as speeding, sudden braking, or driving in the wrong direction. By flexibly adjusting the threshold, the system can adapt to the needs of different situations, optimize detection performance, and provide an efficient and intelligent solution for urban traffic management and safety.

[0090] S6.2: By calculating the anomaly score of each interaction, a sorted list is formed, and the abnormal tracks with high scores are The traffic management department can then focus on high-risk driving behaviors and promptly identify and handle abnormal trajectories, such as speeding, sudden braking, or driving in the wrong direction, thereby improving the safety and stability of the traffic system. This integration process fully utilizes the graph structure information of the traffic network and deep learning technology to provide a powerful and efficient solution for traffic trajectory anomaly detection.

[0091] The second aspect of the present invention relates to a traffic trajectory anomaly detection device based on neighborhood reconstruction and graph contrast learning, characterized in that it includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning of the present invention.

[0092] A computer-readable storage medium of the present invention stores a program thereon, and when the program is executed by a processor, the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning of the claims is implemented.

[0093] The innovation of the present invention is that by combining graph neural network with graph contrast learning technology, the limitations of traditional traffic trajectory anomaly detection methods are broken through, and efficient recognition of complex traffic behavior patterns is achieved. Through deep modeling of the relationship between vehicle nodes and their neighbor nodes, abnormal behaviors at the vehicle level and vehicle interaction level can be accurately captured. At the same time, the present invention introduces the ability to recognize mixed abnormal behavior patterns, so that the system can not only detect a single type of anomaly, but also handle complex patterns composed of multiple abnormal behaviors.

[0094] The working principle of the present invention is (analyzing the reasons for the advantages of the invention); the neighborhood information of traffic trajectory data is reconstructed through a graph neural network (GNN), and the feature vector of the vehicle node and the relationship characteristics with the neighboring nodes are restored. Through multiple GCN layers, the representation of the vehicle nodes is aggregated and updated, so that each node can learn deeper traffic behavior characteristics. In addition, the graph contrast learning technology is adopted to help the model capture abnormal behaviors at the node level and subgraph level through the comparison of positive and negative sample pairs. Finally, through multiple detections and fusion of different anomaly scoring mechanisms, combined with dynamically adjusted thresholds, the accuracy and stability of traffic trajectory anomaly detection are improved. At the same time, the present invention can also identify mixed abnormal behavior patterns, that is, it can simultaneously detect complex abnormal patterns composed of different types of abnormal behaviors (such as speeding, sudden braking, reverse driving, lane deviation, etc.), thereby more accurately reflecting the diverse abnormal behaviors in traffic scenes.

[0095] The advantages of the present invention are:

[0096] 1. High-precision anomaly detection: Combine graph neural networks with graph comparative learning to accurately identify abnormal traffic driving patterns.

[0097] 2. Strong robustness: Adapt to various traffic scenarios and reduce false alarms and missed alarms.

[0098] 3. Flexible threshold adjustment: Use extreme value theory to dynamically adjust the threshold, optimize the detection strategy, and improve system adaptability.

[0099] 4. Efficient data processing: Utilize graph neural networks and deep learning to process large-scale data with good scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 It is a structural diagram of the method of the present invention;

[0101] Figure 2 is a flow chart of the method of the present invention;

[0102] Figure 3 It is a traffic trajectory data construction diagram;

[0103] Figure 4 It is a POT threshold selection diagram of the method of the present invention;

[0104] Figure 5 It is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0105] In order to make the objectives, technical solutions and advantages of the present invention more clear, the specific implementation modes of the present invention will be further described in detail below.

[0106] Example 1

[0107] The present invention provides a method for detecting traffic trajectory anomalies based on neighborhood reconstruction and graph contrast learning. The system flow is as follows: Figure 1 and Figure 2 As shown, the method includes:

[0108] S1: Data cleaning is performed based on the original traffic trajectory data. The steps are as follows:

[0109] S1.1: The original data comes from the Porto dataset, which is widely used in traffic trajectory anomaly detection tasks, especially for analyzing urban traffic patterns. The Porto dataset contains the driving trajectory data of taxis in Porto, including the trajectory records of thousands of taxis and their traffic relationships. Each trajectory record contains the location (latitude and longitude), timestamp, speed, acceleration and other information of the taxi. Each taxi in the data is regarded as a node, and the relationships between taxis (such as proximity, influence, etc.) are represented by edges. The attributes of the data include:

[0110] Table 1

[0111] serial number name Notes 1 Vehicle ID Each taxi in the dataset has a unique identifier 2 Location Each track point contains the current location of the taxi 3 Timestamp Each track point also records the time information 4 speed Vehicle speed 5 Acceleration Vehicle acceleration 6 Road Type Different types of roads may affect the vehicle's driving pattern …… ……

[0112] S1.2: As can be seen from the table, the Porto dataset contains rich traffic trajectory information, including the location information, speed, acceleration, timestamp, etc. of taxis. In order to ensure data quality and availability, data processing is required, such as removing noise points, excluding abnormal trajectories, and correcting erroneous speed or location data. These steps help ensure data integrity, consistency, and accuracy. By cleaning and preprocessing the data, a more reliable traffic trajectory anomaly detection model can be built, and abnormal behaviors in traffic, such as speeding, driving in the wrong direction, and sudden braking, can be analyzed and identified more accurately, thereby improving the safety and efficiency of traffic management and ensuring the normal operation and safety of the traffic system.

[0113] The following methods are used for data cleaning: 1) Remove invalid or abnormal records, such as deleting trajectory points with incomplete location data or beyond a reasonable range, and remove noise points caused by positioning system errors or weak signals to ensure the authenticity of trajectory data. 2) Fill in missing vehicle feature data, such as speed, acceleration, etc., which can be filled using mean interpolation or features based on similar trajectories to ensure the integrity and consistency of each trajectory point. Through these cleaning steps, the high quality and accuracy of the Porto dataset can be guaranteed, providing a reliable data foundation for traffic trajectory anomaly detection.

[0114] S1.3: Construct graph structure data. When processing the Porto dataset, the traffic trajectory data is converted into a graph structure. The trajectory of each taxi is regarded as a node in the graph, and the node features include the vehicle's location information (latitude and longitude), speed, timestamp, etc. The edges in the graph represent the relationship between vehicles, and are usually composed of trajectories traveling on the same road section or adjacent intersections in the same time period. By taking the road network (intersections, road sections) as nodes and constructing edges based on the relationship between vehicles, a traffic flow graph structure is formed. In addition, the weight of the edge can be defined by the distance or travel time between vehicles. This graph structure provides basic data for subsequent anomaly detection and pattern recognition.

[0115] S1.4: For a given undirected graph G = (V, E), where {V 1 , V 2 , V 3 ……V N} represents the set of nodes, N represents the number of nodes, and E represents the set of edges. In addition, the node feature matrix X∈R N×D Represents node feature information, adjacency matrix A∈R N×N Indicates the graph structure information. At the same time, use x i ∈R D Represents node v i The characteristics of d i Represents the degree of each node. For the adjacency matrix A, if A ij =1, it means node v i and v j There is an edge between them, otherwise A ij =0.

[0116] S1.5: Ensure that data is evenly distributed by setting a split ratio of 80% training data and 20% test data. Random sampling is used to allocate samples to training and test sets. The training set is used for model learning and optimization, and the test set is used to evaluate the generalization ability of the model. This step helps to accurately evaluate the model effect and provide a basis for parameter optimization.

[0117] S2: Reconstructing the neighborhood information of graph structured data to capture the similarities and differences in attribute space, such as Figure 3 , the specific steps are as follows:

[0118] S2.1: In traffic trajectory anomaly detection, the neighborhood reconstruction module encodes the node's receptive field through a graph neural network (GNN) and reconstructs the node's neighborhood structure in the reduced-dimensional space. This process not only restores the characteristic attributes of the node, but also reconstructs the connection pattern and attribute relationship between the node and its neighboring nodes, effectively capturing abnormal information in the trajectory network. The method not only uses node representation to reconstruct local neighborhood information, but also introduces comparative learning of global structural information to further improve the modeling ability of complex relationships and attribute spaces between nodes. In this way, the abnormal patterns of nodes can be identified more comprehensively, significantly improving the effect of traffic trajectory network anomaly detection.

[0119] S2.2: Capturing similarities and differences in attribute space. First, the node’s own representation is updated by iteratively aggregating the node’s neighbor information through the GNN Encoder. This aggregation operation can capture the structural information of the node’s local neighborhood and the relationship characteristics with its direct neighbors. The mathematical expression of this process is as follows:

[0120]

[0121] Among them, f i (l) represents the feature vector of node i in layer l, UPDATE represents the operation used to update node features, and the Aggregation function is used to aggregate the information of neighboring nodes. i Represents the set of all neighbor nodes of node i.

[0122] Through a multi-layer perceptron (MLP) i (l+1) Decoded step by step In this way, the original feature vector f is reconstructed i (0) . Then, use l2-loss to calculate With f i (0) The node reconstruction loss is obtained by the difference between:

[0123]

[0124] Reconstructing node degrees using an MLP Its loss function is as follows:

[0125]

[0126] Neighborhood The empirical distribution P i emp Using a multivariate Gaussian distribution To approximate, the mean estimation and covariance matrix estimation formulas are as follows:

[0127]

[0128] Then, from f i (l+1) Constructing an approximate distribution Specifically, using f i (l+1) Generate the mean and covariance matrix of the multivariate Gaussian distribution using the following formula:

[0129]

[0130] From the generated Gaussian distribution k samples are sampled from the dataset, and then these samples are transformed into approximate samples represented by neighbor features through a fully connected neural network (FNN) Finally, use the generated neighbor feature sample To estimate the new mean and covariance matrix:

[0131]

[0132] Based on the given and The KL divergence between these two distributions is used to measure the reconstruction loss of neighbor attribute features:

[0133]

[0134] Where p represents the dimension.

[0135] Finally, the total loss in the neighborhood reconstruction module is as follows:

[0136]

[0137] S3: Subgraph and node sampling to obtain positive and negative sample pairs and extract the potential representation of the subgraph embedding vector. The specific steps are as follows:

[0138] S3.1: First, the node features are initialized, and each node is assigned an initial feature vector. Then, through multiple rounds of convolution operations in the GCN layer, the node features are gradually updated and aggregated, taking into account the relationship between the node and its neighbors. This process enables each node to gradually gather information from surrounding nodes to form a richer representation. It can be expressed by the following formula:

[0139]

[0140] in represents the symmetric normalized adjacency matrix, represents the hidden representation of the lth layer, W (l) Represents weight.

[0141] By extracting the embedding vector of the subgraph through GCN, the model can more accurately understand the features in the graph structure and provide a powerful feature representation for subsequent anomaly detection.

[0142] S3.2: In order to increase the connectivity between nodes in the graph, skip connections are used to fuse the underlying spatial features. The new pooling readout module mainly includes the following steps. Find the indexes of the k largest values:

[0143]

[0144] Then, in the original adjacency matrix Extract the corresponding subgraph adjacency matrix from

[0145]

[0146] Furthermore, element-wise matrix multiplication is performed to obtain the new feature matrix:

[0147]

[0148] Finally, the original structure of the graph is restored through the distribute(·) operation and the final representation of the subgraph z is obtained i , the formula is as follows:

[0149] z i =distribute(0 n×c ,X (l+1) ,idx) (13)

[0150] Accordingly, MLP is used to transform the target node features into the same embedding space as the subgraph, and the final representation of the node is obtained. i , and shares weight W with the previous GCN (l) :

[0151] e i =σ(X (l) W (l) ) (14)

[0152] S4: The graph contrastive learning network captures the similarities and differences of embedding vectors. The specific steps are as follows:

[0153] S4.1: Get the embedding representation of positive and negative sample pairs. The process of graph contrastive learning network capturing the similarity and difference of embedding vectors is mainly achieved by comparing the similarity of positive and negative sample pairs.

[0154] A bilinear model is used to measure the relationship between them, which is calculated by the following formula:

[0155]

[0156] Then, binary cross entropy (BCE) is used to calculate the node-subgraph comparison training loss value:

[0157]

[0158] At the same time, it is also necessary to compare between nodes, which is more conducive to discovering node-level anomalies. Similarly, the new GCN layer and pooling layer are used to obtain the GCN embedding of the target node and another node and the final embedding of the target node. It is also possible to construct sample pairs of positive embedding and negative embedding representing node-node comparison, and the inter-node comparison training loss can be expressed as:

[0159]

[0160] In order to better optimize the embedding representation of subgraphs and the feature recognition ability of nodes, a third comparison mode, subgraph-subgraph comparison, is defined. As one of the more popular comparison losses, InfoNCE loss is defined as:

[0161]

[0162] Generally speaking, minimizing the InfoNCE loss is equivalent to maximizing the lower bound of the mutual information between facing views. Finally, in order to speed up the convergence of the model, the losses during contrastive learning training are aggregated:

[0163]

[0164] Where β∈(0,1) is a parameter used to balance node-level abnormal information, is a parameter used to balance the abnormal information at the subgraph level.

[0165] S5: Define the anomaly scoring function. The specific steps are as follows:

[0166] By minimizing the set objective function, the anomaly score of each node can be calculated. By minimizing the objective function, the anomaly score of each node can be calculated. In the neighborhood reconstruction module, the anomaly score is calculated based on the node self-reconstruction error. The reconstruction error is measured by the difference between the reconstructed feature vector and the original feature vector. The calculation method is:

[0167]

[0168] definition in and Represent the similarity of positive and negative pairs respectively. Based on this rule, the anomaly score of the comparison module can be obtained:

[0169]

[0170] Considering the randomness of a single detection, in order to avoid this situation, multiple detections are performed on each node to obtain the anomaly score of each node, and the average value is taken as the final anomaly score. The above formula can be further rewritten as:

[0171]

[0172] Where R represents the number of anomaly detections.

[0173] S6: Traffic trajectory data anomaly identification application process, such as Figure 4 , the specific steps are as follows:

[0174] S6.1: Take a certain anomaly score as the limit. If it exceeds the limit, it is abnormal, otherwise it is normal, that is, get Threshold:

[0175] Threshold=POT(Score Train ,Score Test ,δ) (25)

[0176] POT is a statistical method based on extreme value theory. It determines the threshold dynamically through the estimated risk value. First, the abnormal scores of all points in the training sample and test sample sequence are calculated, and then the risk value δ∈(0,1) is set to obtain the Threshold.

[0177] S6.2: By calculating the anomaly score of each interaction, a sorted list is formed, and the abnormal tracks with high scores are The traffic trajectory anomaly detection provides an efficient solution by integrating graph structure and deep learning technology.

[0178] The implementation application cases show that the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning proposed in the present invention is effective. Compared with other design methods, the present invention combines graph contrast learning with the neighborhood reconstruction module, applies it to traffic trajectory data, and combines three contrast modes to perform efficient anomaly detection. The neighborhood reconstruction module uses a graph neural network (GNN) to encode the receptive domain of the vehicle trajectory, reconstructs the neighborhood relationship between vehicles in the reduced dimensionality space, and effectively captures the abnormal information of the traffic network in the attribute space. In order to enhance the connectivity between trajectories, the GCN layer and jump connection are used to fuse the underlying features. The model processes the Porto dataset into graph structure data as input and outputs the anomaly score of each vehicle. In addition, the POT method is used to dynamically determine the threshold, and whether the trajectory is abnormal is judged based on whether the anomaly score is greater than the threshold. The experiment used real traffic trajectory data, and the results fully demonstrated the feasibility and superiority of the model.

[0179] Example 2 Figure 5 This embodiment relates to a traffic trajectory anomaly detection device based on neighborhood reconstruction and graph contrast learning, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning of embodiment 1.

[0180] Example 3

[0181] This embodiment involves a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning in the claims is implemented.

[0182] What is described is the specific embodiment of the present invention and the technical principle used. If the changes made according to the concept of the present invention do not exceed the spirit covered by the description and drawings, they should still fall within the scope of protection of the present invention.

Claims

1. A traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning includes the following steps: S1: Select widely used traffic trajectory data, preprocess the raw data, and define the corresponding graph structure data; S2: Reconstructing the neighborhood information of graph structured data to capture the similarities and differences in attribute space; S3: Subgraph and node sampling to obtain positive and negative sample pairs and extract the potential representation of the subgraph embedding vector; S4: Graph contrastive learning networks capture similarities and differences in embedding vectors; S5: define anomaly scoring function; S6: Perform traffic trajectory data anomaly detection.

2. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S1 comprises the following steps: S1.1: Select appropriate traffic trajectory dataset; S1.2: Clean the original data, remove invalid or abnormal records, and fill in the missing node feature data; S1.3: Construct graph structure data, abstract the cleaned Porto traffic trajectory data into graph data; each taxi is represented as a node in the graph, and the node features include the vehicle's location information, speed, and acceleration; construct a graph structure that reflects the urban traffic network, and the edge weight is defined according to the distance between vehicles, travel time, or traffic flow density, so as to more accurately capture the dynamic characteristics of the traffic network; S1.4: For a given undirected graph G = (V, E), where {V1, V2, V3...V N } represents the set of nodes, N represents the number of nodes, and E represents the set of edges; in addition, the node feature matrix X∈R N×D Represents node feature information, adjacency matrix A∈R N ×N Represents graph structure information; at the same time, use x i ∈R D Represents node v i The characteristics of d i Represents the degree of each node; for the adjacency matrix A, if A ij =1, it means node v i and v j There is an edge between them, otherwise A ij =0; S1.5: Divide the processed traffic trajectory data into training samples and test samples; 80% of them are used for training and 20% for testing.

3. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S2 comprises the following steps: S2.1: In reconstructing the neighborhood information of graph structured data, the node's receptive field is encoded through the graph neural network GNN, and the node's neighborhood structure is reconstructed in the dimensionality reduction space; the method not only uses the node representation to reconstruct the local neighborhood information, but also introduces comparative learning of global structural information to further improve the modeling ability of complex relationships between nodes and attribute space; S2.2: In order to capture the similarities and differences in the attribute space, the node’s neighbor information is first iteratively aggregated through the GNN encoder to update the node’s own representation. This aggregation operation can capture the structural information of the node’s local neighborhood and the relationship characteristics with its direct neighbors. The feature information is then decoded by the GNN decoder.

4. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S3 comprises the following steps: S3.1: Extracting the embedding vector of the subgraph is a crucial step in traffic trajectory anomaly detection. The method of the graph convolutional network GCN layer is adopted. Before this, in order to make the obtained potential representation more discriminative, the attributes of the target nodes in the subgraph need to be masked in advance; their hidden layer feature representation can be expressed by the following formula: in represents the symmetric normalized adjacency matrix, represents the hidden representation of the lth layer, W (l) represents weight; By extracting the embedding vector of the subgraph through GCN, the model can more accurately understand the features in the graph structure and provide a powerful feature representation for subsequent anomaly detection; S3.2: In order to increase the connectivity between nodes in the graph, skip connections are used to fuse the underlying spatial features; the new pooling readout module mainly includes the following steps; find the indices of the k largest values: Then, in the original adjacency matrix Extract the corresponding subgraph adjacency matrix from Furthermore, element-wise matrix multiplication is performed to obtain the new feature matrix: Finally, the original structure of the graph is restored through the distribute(·) operation and the final representation of the subgraph z is obtained i , the formula is as follows: z i =distribute(0 n×c ,X (l+1) ,idx) (13) Accordingly, MLP is used to transform the target node features into the same embedding space as the subgraph, and the final representation of the node is obtained. i , and shares weight W with the previous GCN (l) : e i =σ(X (l) W (l) ) (14) 5. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S4 comprises the following steps: S4.1: Based on the above operations, the embedding representation of the positive and negative sample pairs is obtained; the process of the graph contrastive learning network capturing the similarity and difference of the embedding vectors is mainly achieved by comparing the similarity of the positive and negative sample pairs; A bilinear model is used to measure the relationship between them, which is calculated by the following formula: Then, binary cross entropy (BCE) is used to calculate the node-subgraph comparison training loss value: At the same time, it is also necessary to compare between nodes, which is more conducive to discovering node-level anomalies; the inter-node comparison training loss can be expressed as: In order to better optimize the embedding representation of subgraphs and the feature recognition ability of nodes, a third comparison mode, subgraph-subgraph comparison, is defined; as one of the more popular comparison losses, InfoNCE loss is defined as: Generally speaking, minimizing the InfoNCE loss is equivalent to maximizing the lower bound of the mutual information between facing views; finally, in order to speed up the convergence of the model, the losses during contrastive learning training are aggregated: Where β∈(0,1) is a parameter used to balance node-level abnormal information, is a parameter used to balance abnormal information at the subgraph level; By maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs, the graph contrastive learning network optimizes the model parameters through the aggregated loss function to achieve the goal of making similar embedding vectors closer and dissimilar embedding vectors more dispersed.

6. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S5 comprises the following steps: S5: By minimizing the set objective function, the anomaly score of each node can be calculated; By minimizing the objective function, the anomaly score of each node can be calculated In the neighborhood reconstruction module, the anomaly score is calculated based on the node self-reconstruction error. The reconstruction error is measured by the difference between the reconstructed feature vector and the original feature vector. The calculation method is: definition in and Represent the similarity of positive and negative pairs respectively; based on this rule, the anomaly score of the comparison module can be obtained: Considering the randomness of a single detection, in order to avoid this situation, multiple detections are performed on each node to obtain the anomaly score of each node, and the average value is taken as the final anomaly score; the above formula is further rewritten as: Where R represents the number of anomaly detections.

7. The traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as claimed in claim 1, characterized in that: The step S6 comprises the following steps: S6.1: Take a certain anomaly score as the boundary. If it exceeds the boundary, it is abnormal, otherwise it is normal, that is, obtain the threshold Threshold: Threshold=POT(Score Train ,Score Test ,δ) (25) Among them, POT is a statistical method based on extreme value theory. It determines the threshold dynamically through the estimated risk value. First, the abnormal scores of all points in the training sample and test sample sequence are calculated, and then the size of the risk value δ∈(0,1) is set to obtain the Threshold; S6.2: By calculating the anomaly score of each interaction, a sorted list is formed, and the abnormal tracks with high scores are Flagged as a possible anomaly; In this way, traffic management departments can focus on high-risk trajectory behaviors, realize timely identification and processing of abnormal behaviors, and improve traffic safety and efficiency; by integrating graph structures and deep learning technologies, traffic trajectory anomaly detection provides an efficient solution.

8. A traffic trajectory anomaly detection device based on neighborhood reconstruction and graph contrast learning, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the traffic trajectory anomaly detection method based on neighborhood reconstruction and graph contrast learning described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Abnormal traffic flow detection method and system based on dynamic graph

    CN120412286A

  • Wind turbine generator blade edge trajectory abnormal fault diagnosis method and system

    CN120672612A

  • A wind turbine blade edge trajectory abnormality fault diagnosis method and system

    CN120672612B