Vehicle speed anomaly detection method and electronic equipment

By constructing speed feature values ​​and partitioning codes for trajectory data, and combining them with the isolated forest algorithm, the limitations of data distribution assumptions and real-time issues in vehicle speed anomaly detection are resolved, achieving efficient and accurate anomaly detection and reducing false alarm and false negative rates.

CN121838487APending Publication Date: 2026-04-10TUS CLOUD CONTROL (BEIJING) TECH LTD
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TUS CLOUD CONTROL (BEIJING) TECH LTD
Filing Date
2026-01-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies for vehicle speed anomaly detection suffer from limitations in data distribution assumptions, insufficient real-time performance and computational efficiency, limited ability to process high-dimensional and complex data, insufficient generalization ability and robustness, and high data annotation costs.

Method used

The isolated forest algorithm is used to train the speed anomaly detection model. By constructing speed feature values ​​and partitioning codes for trajectory data, the model adapts to different traffic scenarios using geographical features. Data preprocessing, such as missing value imputation and data cleaning, is also performed to reduce dependence on data distribution.

Benefits of technology

It improves the adaptability and accuracy of the model, reduces false alarm and false negative rates, meets the needs of real-time detection, reduces data labeling costs, and adapts to non-normally distributed traffic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838487A_ABST
    Figure CN121838487A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle speed anomaly detection method and electronic equipment. The method comprises the following steps: acquiring trajectory data of a plurality of vehicles; one or more feature values of each piece of trajectory data are constructed, the trajectory data comprise continuous multi-frame vehicle data, the feature values at least comprise speed feature values and partition codes, and the partition codes are codes of areas where geographic positions of vehicles in the trajectory data are located; the characteristic values of the trajectory data with the same partition codes are input into a speed anomaly detection model, abnormal trajectory data output by the speed anomaly detection model are obtained, and the speed anomaly detection model is obtained through training of an isolated forest algorithm. The method can adapt to different traffic scenes and abnormal types, and the false alarm rate and the missing report rate are reduced. In addition, an isolated forest algorithm is used for training a speed anomaly detection model, and the requirement for real-time anomaly detection is met. Meanwhile, the algorithm does not depend on specific distribution of data, and traffic data in non-normal distribution can be processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle-related technologies, and in particular to a method for detecting abnormal vehicle speed, an electronic device, a storage medium, and a computer program product. Background Technology

[0002] With the development of intelligent transportation systems and smart cities, the collection and analysis of vehicle operation data plays an increasingly important role in traffic management. In-depth analysis of real-time vehicle location, speed, acceleration, and other data enables real-time monitoring of traffic conditions, optimization of traffic flow, and early warning of accidents. This is of great significance for improving traffic safety and alleviating traffic congestion. However, with the increase in the number of vehicles and the increasing complexity of the traffic environment, how to efficiently and accurately detect abnormal vehicle speeds has become an urgent problem to be solved. Abnormal vehicle speeds can be caused by a variety of factors, such as abnormal driver behavior, vehicle malfunctions, and changes in road conditions. Timely and accurate detection of these anomalies is crucial for improving traffic safety and optimizing the allocation of traffic resources.

[0003] Meanwhile, real-time dynamic perception of traffic participants is the foundation for ensuring the safe driving of autonomous vehicles and realizing multi-terminal collaborative control. Due to the complexity of technology, cost and scenarios, autonomous vehicles are often applied to specific scenarios, and the perception devices used are not the same. Currently, the more mature devices for detecting abnormal vehicle speeds include onboard equipment, roadside cameras, millimeter-wave radar, and lidar.

[0004] Existing technologies employ various methods and techniques for detecting abnormal vehicle speeds. The main technical solutions are as follows: (1) Threshold-based detection method: Fixed threshold setting: A fixed upper or lower speed limit is set, and when the vehicle speed exceeds this range, it is considered abnormal. For example, the manufacturing method of a car overspeed alarm sets a fixed speed threshold, and an alarm is triggered when the vehicle speed exceeds this threshold. However, although this method is simple to implement, it cannot adapt to different road conditions and traffic situations.

[0005] Dynamic threshold setting: Speed ​​thresholds are dynamically adjusted based on road type, traffic conditions, or time period to improve detection accuracy. For example, a vehicle speed anomaly detection method based on road type uses road information to dynamically adjust the speed threshold.

[0006] (2) Statistical methods: Historical data analysis: Utilizing a large amount of historical vehicle speed data, statistical characteristics such as average speed and standard deviation are calculated. When a vehicle speed deviates from these statistical characteristics within a certain range, it is judged as an anomaly. For example, vehicle speed anomaly detection methods based on big data analysis use historical speed data for statistical analysis.

[0007] Anomaly distribution model: Establish a statistical distribution model of vehicle speed and use a probability density function to detect abnormal speeds. For example, a probabilistic statistical method for vehicle speed anomaly detection constructs a speed distribution model to identify anomalies.

[0008] (3) Methods based on multi-source data fusion: Sensor data fusion: Combining data from multiple sensors such as GPS, accelerometers, and gyroscopes improves the accuracy of speed anomaly detection. For example, a vehicle abnormal behavior detection method based on multi-sensor fusion integrates data from various sensors. Utilizing data from vehicle-to-everything (V2X) and roadside equipment enables real-time monitoring and anomaly detection of vehicle speed.

[0009] (4) Rule-based and expert system-based methods: Rule-based matching: Establish a rule base containing various abnormal situations, and determine anomalies by matching rules based on vehicle speed and other characteristics. The rule-based vehicle speed anomaly detection method either uses a pre-set anomaly rule base for detection or leverages expert knowledge to establish an inference mechanism to identify complex speed anomalies.

[0010] However, the existing technology has the following technical problems: 1. Constraints on assumptions regarding data distribution Traditional statistical methods, such as 3-Sigma and IQR, assume that the data follows a normal distribution or a specific statistical distribution. However, real-world traffic data often does not meet these assumptions, leading to poor detection results.

[0011] 2. Insufficient real-time performance and computational efficiency For example, methods based on multi-source data fusion and anomaly detection methods based on machine learning, such as autoencoders and SVMs, have high computational complexity and long model training and prediction times, which cannot meet the needs of real-time processing.

[0012] 3. Limited ability to process high-dimensional and complex data. Existing methods, such as those based on multi-source data fusion, struggle to fully utilize the spatiotemporal characteristics and vehicle dynamics features of multi-dimensional, high-frequency, and real-time vehicle speed data.

[0013] 4. Insufficient generalization ability and robustness Existing methods, such as dynamic threshold setting methods, suffer from insufficient generalization and are not machine learning methods, so they can only detect specific types. In addition, rule-based and expert system-based methods also suffer from insufficient generalization and robustness, thus they are poorly adaptable to unknown abnormal patterns and diverse traffic scenarios, and are prone to false alarms and false negatives.

[0014] 5. High data annotation costs Existing technologies, such as rule-based and expert system-based methods, are too costly. Supervised learning methods require large amounts of labeled data, which is costly and difficult to obtain in practice. Summary of the Invention

[0015] Therefore, it is necessary to provide a vehicle speed anomaly detection method, electronic device, storage medium, and computer program product to address the technical problems existing in the current technology for vehicle speed anomaly detection.

[0016] This invention provides a method for detecting abnormal vehicle speed, comprising: Acquire trajectory data for multiple vehicles; Construct one or more feature values ​​for each of the trajectory data, the trajectory data including vehicle data of multiple consecutive frames, the feature values ​​including at least speed feature values ​​and partition codes, the partition codes being the codes for the geographical location of the vehicle in the trajectory data; The feature values ​​of the trajectory data with the same partition coding are input into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model, which is trained using the isolated forest algorithm.

[0017] Furthermore, after acquiring the trajectory data of multiple vehicles, the process also includes: For each trajectory data, delete the data that does not meet the standard.

[0018] Furthermore, after acquiring the trajectory data of multiple vehicles, the process also includes: Missing values ​​in the trajectory data are filled in.

[0019] Furthermore, the imputation of missing values ​​in the trajectory data includes: Based on the vehicle data of the preset number of frames preceding the missing value, the vehicle data of the missing frame containing the missing value is calculated as follows: Where S is the predicted vehicle position, S0 is the vehicle position of the closest frame before the missing frame in the trajectory data, v0 is the vehicle speed of the closest frame before the missing frame in the trajectory data, a is the average acceleration of a preset number of frames before the missing value, and t is the time interval between frames in the trajectory data. The predicted vehicle locations are used as vehicle data to fill in the missing frames.

[0020] Furthermore, after acquiring the trajectory data of multiple vehicles, the process also includes: Exclude trajectory data with fewer than a preset frame rate threshold.

[0021] Furthermore, after acquiring the trajectory data of multiple vehicles, the process also includes: The trajectory data is partitioned according to its geographical location.

[0022] Furthermore, the construction of one or more feature values ​​for each of the trajectory data includes: For each of the trajectory data described: Based on multiple frames of vehicle data in the trajectory data, calculate one or more speed feature values ​​that are statistically relevant to the trajectory data; The partition code of the trajectory data is obtained as the partition code of the trajectory data, and the partition code is a one-hot code.

[0023] This invention provides an electronic device, comprising: At least one processor; and, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions that can be executed by at least one of the processors to enable at least one of the processors to perform the vehicle speed anomaly detection method as described above.

[0024] The present invention provides a storage medium that stores computer instructions, which, when executed by a computer, are used to perform all the steps of the vehicle speed anomaly detection method as described above.

[0025] This invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the vehicle speed anomaly detection method as described above.

[0026] This invention constructs feature values ​​from trajectory data, introducing speed feature values ​​and partitioning encoding to fully utilize the characteristics of the data. By incorporating geographical features, the model can adapt to the traffic characteristics of different geographical regions, thus making it more adaptable to different traffic scenarios and anomaly types, reducing false positives and false negatives. Furthermore, the speed anomaly detection model is trained using the Isolation Forest algorithm. The Isolation Forest algorithm is highly efficient in processing high-dimensional, large-scale data, making it suitable for real-time anomaly detection needs. Simultaneously, the Isolation Forest algorithm is unsupervised learning, requiring no large amount of labeled data, reducing data annotation costs. Therefore, the algorithm does not depend on a specific data distribution and can handle non-normally distributed traffic data. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating the process of a vehicle speed anomaly detection method according to an embodiment of the present invention. Figure 2This is a flowchart illustrating a vehicle speed anomaly detection method according to another embodiment of the present invention. Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to the present invention. Detailed Implementation

[0028] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings. Identical components are indicated by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "up," and "down" used in the following description refer to directions in the accompanying drawings, while the terms "inner" and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.

[0029] like Figure 1 The diagram shown is a flowchart of a vehicle speed anomaly detection method according to an embodiment of the present invention, including: Step S101: Obtain trajectory data for multiple vehicles; Step S102: Construct one or more feature values ​​for each of the trajectory data, the trajectory data including vehicle data of multiple consecutive frames, the feature values ​​including at least speed feature values ​​and partition codes, the partition codes being the codes of the geographical location of the vehicle in the trajectory data; Step S103: Input the feature values ​​of the trajectory data with the same partition code into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model. The speed anomaly detection model is trained using the isolated forest algorithm.

[0030] Specifically, the present invention can be applied to electronic devices with processing capabilities, such as vehicle controllers or cloud servers.

[0031] First, execute step S101 to obtain trajectory data for multiple vehicles.

[0032] Specifically, when operating in a vehicle, the vehicle needs to collect trajectory data from other vehicles as well as its own to identify abnormal trajectory data. When operating on a cloud server, the vehicle sends trajectory data to the cloud, allowing the cloud server to collect a large amount of trajectory data from various vehicles to identify abnormal trajectory data.

[0033] Then, step S102 is performed to construct one or more feature values ​​for each of the trajectory data, the trajectory data including vehicle data of multiple consecutive frames, the feature values ​​including at least speed feature values ​​and partition codes, the partition codes being the codes for the geographical location of the vehicle in the trajectory data.

[0034] Specifically, each vehicle provides one or more trajectory data sets. The trajectory data consists of multiple consecutive frames of vehicle data, such as more than 50 consecutive frames. Vehicle data includes, but is not limited to: speed, acceleration, longitude, latitude, timestamp, heading angle, data source type, unique identifier (UUID), and partition code.

[0035] Then, feature values ​​are calculated for each frame of data related to the trajectory. These feature values ​​include at least a speed feature value and a partition code, where the partition code is the code for the geographical location of the vehicle in the trajectory data. Specifically, the speed feature value is a speed-related feature value calculated based on the vehicle data within the trajectory data.

[0036] Velocity characteristic values ​​can be statistical features of velocity. Velocity characteristic values ​​include, but are not limited to, maximum velocity, minimum velocity, average velocity, maximum acceleration, minimum acceleration, and average acceleration.

[0037] The data is coded for different geographical regions. Since traffic characteristics can vary significantly across different geographical regions, they need to be processed separately. Therefore, the data is partitioned based on the geographical location (latitude and longitude) of the vehicles, and each partition is coded. The preferred number of partitions is 24.

[0038] Finally, step S103 is executed, in which the feature values ​​of the trajectory data with the same partition code are input into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model, which is trained using the isolated forest algorithm.

[0039] Specifically, the feature values ​​of the trajectory data with the same partition code are input into the speed anomaly detection model, and an independent speed anomaly detection model is used for each partition code to predict abnormal trajectory data. The speed anomaly detection model is trained using the Isolation Forest algorithm. The Isolation Forest algorithm is a machine learning algorithm for anomaly detection. It constructs a tree based on input features and calculates anomaly scores for nodes to detect anomalies. In this embodiment, one or more feature values ​​of each trajectory data are input as a set of feature values ​​into the speed anomaly detection model. Each set of feature values ​​for each trajectory data serves as a node in the Isolation Forest algorithm. Anomaly nodes are calculated using the Isolation Forest algorithm, and the trajectory data corresponding to these anomaly nodes are the abnormal trajectory data. Simultaneously, the Isolation Forest algorithm can also output anomaly scores for the abnormal trajectory data as anomaly values.

[0040] The main parameters of the Isolation Forest algorithm typically include the number of trees (n_estimators), the sampling size (max_samples), the maximum tree depth (max_features), and the contamination factor. The selection of these key parameters has a significant impact on the algorithm's performance and efficiency when implementing the Isolation Forest algorithm.

[0041] `n_estimators` (number of trees): This parameter defines the total number of isolation trees that make up the forest. A larger number of trees can improve the stability and accuracy of the model, but it will also increase the computational cost. For small to medium-sized datasets, a range of 100 to 200 trees is generally recommended.

[0042] `max_samples` (maximum sample size): This parameter determines the number of data points used to build each isolation tree. The performance of isolation forests is positively correlated with the selected sample size, but a larger sample size may lead to a decrease in computational efficiency. Therefore, a balance needs to be found between algorithm performance and computational efficiency.

[0043] Contamination (Outlier Ratio): This parameter estimates the proportion of outliers in the dataset. This ratio can be adjusted depending on the characteristics of the dataset and the specific anomaly detection objective. The default value is usually set to 'auto', allowing the algorithm to estimate it automatically based on the data.

[0044] Randomness control (random_state): To ensure the reproducibility of results, a fixed random seed is usually required. This is especially important when evaluating and comparing algorithm performance to ensure consistency of experimental conditions.

[0045] When applying the Isolation Forest algorithm to detect anomalies in a dataset, the choice of parameters is not only affected by the size and features of the dataset, but also by the expected anomaly detection target.

[0046] Outlier identification The Isolation Forest algorithm is used to train and predict data, identifying outliers. When the algorithm performs anomaly detection on the data in the uuid_stats.csv dataset, the main output fields reflecting the algorithm results include the following two types: (1) Anomaly score: Isolation forest provides a score for each data point, reflecting the degree of anomaly of the data point. The lower the score of the vehicle track, the more likely the vehicle track is to be an anomaly.

[0047] (2) Abnormal / Non-abnormal labels: Data points can be labeled as abnormal or non-abnormal based on their scores. Typically, data points with scores below a certain threshold are labeled as abnormal.

[0048] Model evaluation methods Data visualization For raw data, histograms and box plots can be used to view the distribution of individual features, or scatter plots can be used to view the relationships between features.

[0049] Since the data model runs on a Linux server without a graphical interface, directly displaying dynamic charts is not feasible. Visualization can be achieved by saving the corresponding charts as image files and exporting them. The visualization uses matplotlib's `savefig` method.

[0050] For visualized data, the analysis process requires further review of data points marked as outliers and evaluation of the overall model performance. The specific steps are as follows: (1) Examine abnormal data points: Examine the data points marked as anomalous in detail, analyze their characteristics and contextual information, and try to understand why they were judged as anomalous. Compare the characteristics of anomalous data points with normal data points to look for potential patterns or causes of anomalous behavior.

[0051] (2) Evaluate model performance: By comparing the outliers labeled by the algorithm with known outliers or expert annotations, the model's accuracy, recall, and F1 score are evaluated. Analyzing false positives (normal data mislabeled as outliers) and false negatives (outlier data not correctly labeled) helps identify potential areas for model improvement.

[0052] (3) Adjust the threshold and parameters: Adjust the anomaly score threshold based on model performance and business requirements to balance false positives and false negatives. Consider adjusting model parameters or preprocessing steps to improve model accuracy and reliability.

[0053] Model evaluation metrics For model evaluation, different evaluation metrics may be needed for each speed anomaly event. For example, accuracy might be an important metric for speeding detection, while recall might be more important for emergency braking to avoid excessive false negatives. Therefore, this paper will use multiple evaluation metrics (such as accuracy, recall, F1 score, etc.) to comprehensively evaluate the model's performance. The model evaluation is conducted using the following methods: (1) Preparation for verification work First, a set of known abnormal and non-abnormal vehicle trajectory data is acquired or created as a baseline for the model results data. This process is usually accomplished through manual labeling by experts, recording historical anomaly reports, or using simulated data. This paper, however, completes the addition of simulated abnormal data during the data preprocessing stage and obtains abnormal parking, speeding, emergency braking, and abnormal low-speed warning data output by algorithms from other scenarios from real log records.

[0054] (2) Define verification metrics The evaluation process can employ a binary classification approach, dividing vehicle trajectory data containing speed into normal data (represented by 0) and abnormal data (represented by 1). The detected data and the true values ​​can form a binary classification table as follows: Binary Classification Results Comparison Table

[0055] From the table above, we can see that T p (True Positive) indicates that the model's detection results are consistent with the actual vehicle trajectory, meaning there are no anomalies; F p (False Positive) indicates that the model detects an anomaly in the vehicle trajectory, but the actual vehicle operation is normal; F n (False Negative) indicates that the model detects normal vehicle trajectories, but the actual vehicle trajectories should be abnormal data; T n (True Negative) indicates that the vehicle trajectory detected by the model is abnormal data, and the actual vehicle operation value is indeed abnormal.

[0056] The four indicators mentioned above work together on the confusion matrix, and it can be seen that when the true class T... p and true negative class T n The larger the value of F, the more accurate the model's detection, because in this case, the model identifies both normal and abnormal data as the same as the true value. Conversely, a smaller value indicates a false negative class F. n and false positive class F p The larger the value, the less accurate the model becomes, and these values ​​correspond to cases of missed reports and false alarms, respectively.

[0057] By calculating the above four indicators, we can obtain the accuracy, recall, precision, F1 score, and other indicators of the model's detection performance. These indicators can then be used for quantitative analysis.

[0058] Accuracy represents the proportion of tracks where the model correctly identifies normal and abnormal speed behavior. In the context of vehicle speed anomaly detection, this means the proportion of data points where the algorithm correctly distinguishes between normal and abnormal speeds, and its calculation formula is as follows:

[0059] Recall reflects the proportion of speeding or abnormal deceleration events captured by the algorithm to all actual such events. A high recall indicates that the model can miss fewer true abnormal speed behaviors, and is a key metric focused on in this study. Its calculation formula is as follows:

[0060] Precision indicates the proportion of data that, out of events judged as speeding or abnormal deceleration, are actually abnormal. High precision indicates that the model is more accurate in labeling anomalies, i.e., a lower false positive rate. Its calculation formula is as follows:

[0061] The F1 score, through harmonic averaging, ensures a balanced approach to precision and recall. It avoids prioritizing high recall at the expense of precision, which could lead to increased false positives, or prioritizing precision at the expense of recall, which could result in severe false negatives (Bekkar et al., 2013). Therefore, the F1 score provides a balanced perspective on these two aspects, allowing for a more comprehensive understanding of the model's performance in accurately identifying outliers. Its calculation formula is as follows:

[0062] By calculating and analyzing these indicators, we can gain a deeper understanding of the overall performance of the model and guide subsequent optimization efforts.

[0063] (3) Analyze the causes of specific flight path locations. For non-anomaly data points that are falsely reported, analyze their characteristics and the model's judgment logic to identify the reasons for the false reports. For anomaly data points that are missed, analyze whether the model failed to correctly identify them because their characteristics do not conform to common anomaly patterns or the anomaly level is not significant enough. Based on the results of the validation analysis, adjust the parameter settings of the Isolation Forest algorithm, reset the anomaly score threshold, or the number of trees, etc., to improve the model's accuracy.

[0064] (4) Re-evaluation and selection of features Record the results of the current features and parameters of the algorithm, and consider whether it is necessary to optimize the data preprocessing steps, introduce new features, eliminate noisy data, or further clean the data.

[0065] By verifying the accuracy of outliers through the above process, the effectiveness and reliability of the Isolation Forest algorithm in vehicle speed anomaly detection applications can be ensured. During the algorithm parameter debugging process, parameters suitable for engineering applications are gradually discovered, providing practical guidance for continuous model improvement.

[0066] In some embodiments, inputting the feature values ​​of the trajectory data having the same partition coding into the velocity anomaly detection model includes: The feature values ​​of the trajectory data that have the same partition coding and are in the same time period are input into the velocity anomaly detection model.

[0067] Specifically, a day can be divided into multiple time periods. The feature values ​​of trajectory data with the same partition code and falling within the same time period are input into the speed anomaly detection model. The resulting abnormal trajectory data is then specific to that time period. For example, a day can be divided into daytime (06:00-22:00) and nighttime (22:00-06:00). The actual number of segments can be increased; the more detailed the time feature segmentation, the more specific the resulting data will be to a particular time period. This fully utilizes the spatiotemporal characteristics of the data. By incorporating geographical and temporal features, the model can adapt to the traffic characteristics of different geographical regions and time periods.

[0068] This invention constructs feature values ​​from trajectory data, introducing speed feature values ​​and partitioning encoding to fully utilize the characteristics of the data. By incorporating geographical features, the model can adapt to the traffic characteristics of different geographical regions, thus making it more adaptable to different traffic scenarios and anomaly types, reducing false positives and false negatives. Furthermore, the speed anomaly detection model is trained using the Isolation Forest algorithm. The Isolation Forest algorithm is highly efficient in processing high-dimensional, large-scale data, making it suitable for real-time anomaly detection needs. Simultaneously, the Isolation Forest algorithm is unsupervised learning, requiring no large amount of labeled data, reducing data annotation costs. Therefore, the algorithm does not depend on a specific data distribution and can handle non-normally distributed traffic data.

[0069] like Figure 2 The diagram shown is a flowchart of a vehicle speed anomaly detection method according to another embodiment of the present invention, including: Step S201: Obtain trajectory data for multiple vehicles.

[0070] Step S202: For each trajectory data, delete the data that does not meet the standard.

[0071] Step S203: Fill in the missing values ​​in the trajectory data.

[0072] In one embodiment, imputing missing values ​​in the trajectory data includes: Based on the vehicle data of the preset number of frames preceding the missing value, the vehicle data of the missing frame containing the missing value is calculated as follows: Where S is the predicted vehicle position, S0 is the vehicle position of the closest frame before the missing frame in the trajectory data, v0 is the vehicle speed of the closest frame before the missing frame in the trajectory data, a is the average acceleration of a preset number of frames before the missing value, and t is the time interval between frames in the trajectory data. The predicted vehicle locations are used as vehicle data to fill in the missing frames.

[0073] Step S204: Exclude trajectory data with fewer than a preset frame count threshold.

[0074] Step S205: Divide the trajectory data into partitions based on the geographical location of the trajectory data.

[0075] Step S206, for each of the trajectory data: Based on multiple frames of vehicle data in the trajectory data, calculate one or more speed feature values ​​that are statistically relevant to the trajectory data; The partition code of the trajectory data is obtained as the partition code of the trajectory data, and the partition code is a one-hot code.

[0076] Step S207: Input the feature values ​​of the trajectory data with the same partition code into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model. The speed anomaly detection model is trained using the isolated forest algorithm.

[0077] This embodiment proposes a vehicle speed anomaly detection method based on the isolated forest algorithm, which specifically includes the following steps: Preprocessing, including data cleaning based on vehicle dynamics and statistical characteristics: A series of data cleaning methods were designed to address the characteristics of vehicle data, including handling data frequency and latency issues, missing value imputation, data filtering, and partitioning. This ensures data quality and consistency.

[0078] Feature engineering involves selecting and constructing features from the data to extract key vehicle dynamics and statistical features, such as speed, acceleration, location, and time. One-hot encoding is used for geographic regions to avoid the influence of numerical values ​​on the model.

[0079] Model Training and Evaluation: An anomaly detection model is constructed using the Isolation Forest algorithm. The model is trained and optimized, and its performance is comprehensively evaluated using various evaluation metrics and visualization methods.

[0080] Real-time anomaly detection and application: Deploy the model in the actual system to perform real-time anomaly detection on vehicle speed, and provide services such as accident warning and traffic management.

[0081] Specifically, step S201 is executed first to obtain trajectory data of multiple vehicles.

[0082] Specifically, when operating in a vehicle, the vehicle needs to collect trajectory data from other vehicles as well as its own to identify abnormal trajectory data. When operating on a cloud server, the vehicle sends trajectory data to the cloud, allowing the cloud server to collect a large amount of trajectory data from various vehicles to identify abnormal trajectory data.

[0083] Then, steps S202 to S205 are performed to preprocess the data. Wherein: In step S202, for each trajectory data, delete the data that does not meet the standard.

[0084] Specifically, the reporting frequency and latency of vehicle data may be abnormal, such as discontinuous or deviating from the normal trajectory. Therefore, monitoring the reporting frequency and latency of data and deleting non-compliant data addresses these issues. Non-compliant data is defined as data with a latency greater than 1 second for 5 consecutive frames, or data reporting frequencies higher than 20Hz or lower than 5Hz.

[0085] Then, step S203 is performed to fill in the missing values ​​in the trajectory data.

[0086] Specifically, the data may contain missing values, affecting model training and prediction. Vehicle data consists of multiple consecutive frames; therefore, there should be a set of vehicle data at each preset time interval. If a value is missing in a certain time interval, it is considered a missing value. The speed and position change trends of vehicle data based on the preset number of frames preceding the missing frame are used to fill in the missing values ​​through a prediction model. The prediction model is a kinematic model, serving as a pre-module of the Isolation Forest algorithm.

[0087] In one embodiment, imputing missing values ​​in the trajectory data includes: Based on the vehicle data of the preset number of frames preceding the missing value, the vehicle data of the missing frame containing the missing value is calculated as follows: Where S is the predicted vehicle position, S0 is the vehicle position of the closest frame before the missing frame in the trajectory data, v0 is the vehicle speed of the closest frame before the missing frame in the trajectory data, a is the average acceleration of a preset number of frames before the missing value, and t is the time interval between frames in the trajectory data. The predicted vehicle locations are used as vehicle data to fill in the missing frames.

[0088] Specifically, a predictive model based on the speed and position change trends of vehicle data from a pre-set number of frames can be used to fill in the gaps, employing kinematic formulas: Where S is the predicted vehicle position, S0 is the vehicle position of the closest frame before the missing frame in the trajectory data, v0 is the vehicle speed of the closest frame before the missing frame in the trajectory data, a is the average acceleration of a preset number of frames before the missing value, and t is the time interval between frames in the trajectory data.

[0089] Preferably, the preset frame count is 10 frames, then S is the predicted position of the 11th frame. S0 is the position of the 10th frame, v0 is the velocity of the 10th frame, a is the average acceleration of the previous 10 frames, and t is the time interval between two frames, preferably 100ms (10Hz). In practice, it can also be dynamically configured according to the upload frequency of the connected vehicle.

[0090] The missing frame is filled with the vehicle's position to complete the trajectory. Simultaneously, the vehicle's acceleration is assumed to remain constant, meaning the acceleration captured in the filler frame is unchanged. The speed is then calculated based on the previous frame: vehicle speed v = v0 + at.

[0091] Then, step S204 is executed to exclude trajectory data with fewer than a preset frame number threshold.

[0092] Specifically, if the trajectory data is too short, it may not provide enough information. Therefore, data filtering is performed to exclude vehicle trajectory data with sampling points below a preset frame rate threshold. Preferably, vehicle trajectory data with fewer than 50 frames are excluded.

[0093] Then, step S205 is executed to partition the trajectory data according to the geographical location of the trajectory data.

[0094] Specifically, traffic characteristics may vary significantly across different geographical regions, requiring separate processing. Therefore, the trajectory data is partitioned according to its geographical location, i.e., the vehicle's geographical location (latitude and longitude). Preferably, the number of partitions is 24. For cross-regional cases, data from adjacent regions can be transferred to the same region for processing; that is, if a vehicle crosses regions, its data remains continuous.

[0095] After completing the data preprocessing, step S206 is executed for each of the trajectory data: Based on multiple frames of vehicle data in the trajectory data, calculate one or more speed feature values ​​that are statistically relevant to the trajectory data; The partition code of the trajectory data is obtained as the partition code of the trajectory data, and the partition code is a one-hot code.

[0096] Step S206 constructs features for the trajectory data.

[0097] The trajectory data includes raw features such as speed, acceleration, longitude, latitude, timestamp, heading angle, data source type, unique identifier (UUID), and partition code.

[0098] Based on the original features of trajectory data, Statistical characteristics: maximum speed, minimum speed, average speed, maximum acceleration, minimum acceleration, average acceleration, etc.

[0099] Meanwhile, for partition coding, since partition coding is a categorical variable, one-hot encoding is used to avoid introducing numerical bias by using it directly.

[0100] Table 1. Example of a zoned one-hot coding table

[0101] By converting the partition encoding to binary representation, five binary partition fields are constructed to avoid the dimension explosion problem.

[0102] Finally, regarding the feature standardization issue, since the Isolation Forest algorithm is insensitive to data scale, there is no need for normalization or standardization processing, and the original values ​​of the features are preserved.

[0103] Finally, step S207 is executed, in which the feature values ​​of the trajectory data with the same partition code are input into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model, which is trained using the isolated forest algorithm.

[0104] Specifically, the first step is to construct a speed anomaly detection model.

[0105] Model building 1. Algorithm selection: Isolation Forest algorithm.

[0106] 2. Parameter settings: n_estimators (number of trees): set to 100.

[0107] max_samples (maximum sample size): Set to "auto" to use all samples.

[0108] Contamination (abnormality rate): Adjusted according to the actual data.

[0109] 3. Training data: Vehicle data that has been cleaned and feature-engineered, including simulated outlier data.

[0110] 4. Training process: The model learns normal and abnormal patterns in the data by constructing a series of randomly split decision trees.

[0111] Model Evaluation Evaluation indicators: Accuracy: The proportion of samples that are correctly predicted out of the total sample.

[0112] Recall: The proportion of correctly identified anomalous samples out of the actual anomalous samples.

[0113] F1 Score: The harmonic mean of precision and recall.

[0114] Data visualization: Use histograms, box plots, scatter plots, and other methods to display data distribution and model results.

[0115] Results analysis: Initial model results: Without incorporating geographic partitioning and temporal features into the model, the model is unable to effectively detect anomalies in certain partitions.

[0116] Optimized model results: After incorporating geographical partitions, the model can more accurately capture the driving patterns and abnormal behaviors of each partition. Furthermore, time features can be incorporated to make the model even more accurate.

[0117] Impact of acceleration features: After introducing acceleration features, the anomaly detection performance of the model is significantly improved, especially in urban roads where vehicle acceleration changes frequently and abnormal behavior is more easily captured.

[0118] The trained speed anomaly detection model can be deployed on a cloud server or on a vehicle. By inputting the feature values ​​of the trajectory data with the same partition coding into the speed anomaly detection model, the abnormal trajectory data output by the speed anomaly detection model can be obtained.

[0119] Vehicles corresponding to abnormal trajectory data are considered abnormal vehicles. Furthermore, the speed anomaly detection model using the Isolation Forest algorithm can output anomaly scores as outlier values ​​for the abnormal trajectory data. Abnormal trajectory data can be used to provide services such as accident warnings and traffic management.

[0120] This embodiment improves the accuracy and robustness of anomaly detection. Data cleaning based on vehicle dynamics and statistical features enhances data quality and reduces the impact of noise on the model. In complex and ever-changing traffic environments, it accurately and promptly detects vehicle speed anomalies, reducing false positives and false negatives. Simultaneously, it utilizes a rich feature set, incorporating multi-dimensional features such as speed, acceleration, geographic location, and time, along with partitioned one-hot encoding, fully leveraging the spatiotemporal characteristics of the data. By introducing acceleration features and geographic partitioning, the model is more adaptable to different traffic scenarios and anomaly types, further reducing false positives and false negatives. Furthermore, this embodiment meets the requirements of real-time processing. The Isolation Forest algorithm is highly efficient in processing high-dimensional, large-scale data, making it suitable for real-time anomaly detection. By appropriately setting model parameters, it balances model performance and computational efficiency. This embodiment designs an efficient algorithm and model capable of processing massive, high-dimensional vehicle speed data, achieving real-time anomaly detection. This embodiment reduces reliance on data distribution assumptions and labeled data. The Isolation Forest algorithm is an unsupervised learning algorithm, requiring minimal labeled data and reducing data annotation costs. By employing unsupervised or weakly supervised machine learning methods, this embodiment avoids strict assumptions about data distribution and reduces the need for large amounts of labeled data. This embodiment is highly adaptable; the algorithm does not depend on a specific data distribution and can handle non-normally distributed traffic data. This embodiment fully utilizes vehicle dynamics and statistical features, enhancing the model's anomaly detection capabilities through in-depth mining of these features. This embodiment enhances the model's generalization and adaptability by adding various types of simulated anomaly data, enabling the model to adapt to different traffic scenarios and anomaly types, improving its ability to detect unknown anomaly patterns. The inclusion of geographical and temporal features allows the model to adapt to traffic characteristics in different geographical regions and time periods. Finally, this embodiment is easy to deploy and apply; the model structure is simple and easy to deploy in practical systems, supporting real-time monitoring and early warning. This embodiment provides an efficient and accurate anomaly detection solution for intelligent transportation systems, applicable to accident early warning, traffic management, navigation services, etc., and has engineering implementation and commercial application value.

[0121] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0122] like Figure 3 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 301; and, A memory 302 communicatively connected to at least one of the processors 301; wherein, The memory 302 stores instructions that can be executed by at least one of the processors to enable the at least one of the processors to perform the vehicle speed anomaly detection method as described above.

[0123] Figure 3 Take processor 301 as an example.

[0124] The electronic device may also include an input device 303 and a display device 304.

[0125] The processor 301, memory 302, input device 303 and display device 304 can be connected by a bus or other means. The figure shows an example of connection by bus.

[0126] The memory 302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the vehicle speed anomaly detection method in the embodiments of this application, for example, Figure 1 , Figure 2 The method flow is shown. The processor 301 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 302, thereby realizing the vehicle speed anomaly detection method in the above embodiment.

[0127] The memory 302 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the vehicle speed anomaly detection method. Furthermore, the memory 302 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 302 may optionally include memory remotely located relative to the processor 301, and these remote memories may be connected via a network to the apparatus performing the vehicle speed anomaly detection method. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0128] The input device 303 can receive user clicks and generate signal inputs related to user settings and function control of the vehicle speed anomaly detection method. The display device 304 may include a display screen or other display equipment.

[0129] When one or more modules are stored in the memory 302, and are run by one or more processors 301, the vehicle speed anomaly detection method in any of the above method embodiments is executed.

[0130] This invention constructs feature values ​​from trajectory data, introducing speed feature values ​​and partitioning encoding to fully utilize the characteristics of the data. By incorporating geographical features, the model can adapt to the traffic characteristics of different geographical regions, thus making it more adaptable to different traffic scenarios and anomaly types, reducing false positives and false negatives. Furthermore, the speed anomaly detection model is trained using the Isolation Forest algorithm. The Isolation Forest algorithm is highly efficient in processing high-dimensional, large-scale data, making it suitable for real-time anomaly detection needs. Simultaneously, the Isolation Forest algorithm is unsupervised learning, requiring no large amount of labeled data, reducing data annotation costs. Therefore, the algorithm does not depend on a specific data distribution and can handle non-normally distributed traffic data.

[0131] One embodiment of the present invention provides a storage medium that stores computer instructions, which, when executed by a computer, are used to perform all the steps of the vehicle speed anomaly detection method described above.

[0132] In the context of this disclosure, a storage medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The storage medium can be a machine-readable signal medium or a machine-readable storage medium. Optionally, the storage medium can be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), compact disc ROM (CD-ROM), magnetic tape, floppy disk, and optical data storage device.

[0133] One embodiment of the present invention provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements the vehicle speed anomaly detection method as described above.

[0134] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for detecting abnormal vehicle speed, characterized in that, include: Acquire trajectory data for multiple vehicles; Construct one or more feature values ​​for each of the trajectory data, the trajectory data including vehicle data of multiple consecutive frames, the feature values ​​including at least speed feature values ​​and partition codes, the partition codes being the codes for the geographical location of the vehicle in the trajectory data; The feature values ​​of the trajectory data with the same partition coding are input into the speed anomaly detection model to obtain the abnormal trajectory data output by the speed anomaly detection model, which is trained using the isolated forest algorithm.

2. The vehicle speed anomaly detection method according to claim 1, characterized in that, After acquiring the trajectory data of multiple vehicles, the process also includes: For each trajectory data, delete the data that does not meet the standard.

3. The vehicle speed anomaly detection method according to claim 1, characterized in that, After acquiring the trajectory data of multiple vehicles, the process also includes: Missing values ​​in the trajectory data are filled in.

4. The vehicle speed anomaly detection method according to claim 3, characterized in that, The process of filling in missing values ​​in the trajectory data includes: Based on the vehicle data of the preset number of frames preceding the missing value, the vehicle data of the missing frame containing the missing value is calculated as follows: Where S is the predicted vehicle position, S0 is the vehicle position of the closest frame before the missing frame in the trajectory data, v0 is the vehicle speed of the closest frame before the missing frame in the trajectory data, a is the average acceleration of a preset number of frames before the missing value, and t is the time interval between frames in the trajectory data. The predicted vehicle locations are used as vehicle data to fill in the missing frames.

5. The vehicle speed anomaly detection method according to claim 1, characterized in that, After acquiring the trajectory data of multiple vehicles, the process also includes: Exclude trajectory data with fewer than a preset frame rate threshold.

6. The vehicle speed anomaly detection method according to claim 1, characterized in that, After acquiring the trajectory data of multiple vehicles, the process also includes: The trajectory data is partitioned according to its geographical location.

7. The vehicle speed anomaly detection method according to any one of claims 1 to 6, characterized in that, The construction of one or more feature values ​​for each of the trajectory data includes: For each of the trajectory data described: Based on multiple frames of vehicle data in the trajectory data, calculate one or more speed feature values ​​that are statistically relevant to the trajectory data; The partition code of the trajectory data is obtained as the partition code of the trajectory data, and the partition code is a one-hot code.

8. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to at least one of the processors; wherein, The memory stores instructions executable by at least one of the processors, which, when executed by at least one of the processors, enable the at least one of the processors to perform the vehicle speed anomaly detection method as described in any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium stores computer instructions, which, when executed by the computer, are used to perform all the steps of the vehicle speed anomaly detection method as described in any one of claims 1 to 7.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the vehicle speed anomaly detection method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Driver state evaluation method and system, and car-road cooperation system

    CN122201002A

  • Driver state evaluation method and system, and car-road cooperation system

    CN122201002B