A Semantics-Aware Abnormal Detection Method for Trajectory Data
Through editing distance and adaptive window technology, combined with the sampling interval and motion speed of trajectory data, the problem of ignoring semantic features in the existing methods is solved, and the trajectory anomaly detection under multi-dimensional semantic constraints is realized, which improves detection accuracy and efficiency.
Patent Information
- Application Number
- CN202210850711.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-19
AI Technical Summary
The existing trajectory data anomaly detection methods mainly consider spatiotemporal features, ignore the multidimensional semantic features of the trajectory, and cannot realize abnormal trajectory detection under semantic feature constraints.
The similarity between cells is calculated by editing distance, combining the sampling interval and average motion speed of track points to determine the grid size, and filtering historical track data through an adaptive window, generating a new working data set, calculating support and outliers, and accurately detecting abnormal track fragments.
Trajectory anomaly detection under multidimensional semantic constraints is realized, which avoids the problems of insufficient detection accuracy and excessive calculation volume caused by improper grid size, and can accurately identify and sort abnormal trajectories.
Smart Images

Figure CN115169476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of anomaly detection methods for trajectory data, in particular to an anomaly detection method for trajectory data considering semantic information. Background Art
[0002] Trajectory data is a time-series data stream generated by various positioning devices, which describes the position changes of moving objects within a certain time range and reflects the activity rules and behavior patterns of the objects. Since anomalies are not equivalent to noise, anomalies may arise from different mechanisms, and the occurrence of anomalies may be accompanied by interesting phenomena. Therefore, the anomaly detection of trajectory data has important research value for large-scale trajectory data mining and knowledge discovery. Currently, the anomaly detection for trajectory data can be classified into the following four categories:
[0003] (1) Classification-based detection methods
[0004] The classification-based method establishes a classification model for the data set according to the labeled data set, and then determines the predetermined category to which the target data belongs through the classification model. The detection process usually consists of two stages. One is to establish a classification model through the labeled data set, which is collectively called the training stage. The other is to divide the target data into the category to which it belongs according to the classification model, and this stage is called the testing stage. The classification-based anomaly detection technology can usually obtain better detection accuracy than unsupervised technology. However, attaching correct labels to the data requires expert manual annotation, and obtaining labeled data that can accurately represent the behaviors of all classes (including normal and abnormal) involves high computational overhead.
[0005] (2) Historical similarity-based detection methods
[0006] The historical similarity-based detection method establishes a global feature model for anomaly detection according to the collected trajectory data, and then uses the global feature model to identify the trajectory data to be detected, and marks the trajectory data that is different from the global feature model as abnormal. Since this method uses a large amount of historical trajectory data, it usually does not consider the time-varying and concept drift of the trajectory data. When the historical data reaches a certain amount, the global feature model established based on it has high accuracy. It has rich applications and good performance in ship anomaly detection and road network traffic. However, since this type of method only relies on historical data for global feature model modeling, it cannot adapt to the time-varying characteristics of trajectory data.
[0007] (3) Clustering-based detection methods
[0008] The clustering-based method can be regarded as an unsupervised learning method. By relying on various features of trajectory data, the trajectory data is divided into different clusters in a clustering manner, maximizing the differences between clusters and minimizing the differences within clusters. The data that does not belong to any cluster or the data with a quantity less than a certain cluster is output as abnormal data. The anomaly detection using the clustering method is an effective method, but there are also two problems: 1) It is an unsupervised learning method and highly depends on the clustering results. 2) Although the abnormal data only accounts for a very small part of the whole, a large amount of normal data still needs to be processed first, resulting in a large computational overhead.
[0009] (4) Method based on grid division
[0010] The method based on grid division is to divide the area into different grids, and then transform the anomaly detection problem into the anomaly recognition of grid sequences. In the urban road network, due to the limitations of road width, speed, etc., the trajectories are highly regular. Therefore, the grid-based algorithm is effective. However, for free and unconstrained spaces, it is difficult to divide the grids, and the trajectory rules are sparse, so this method is not applicable.
[0011] Based on the above research, there have also emerged anomaly trajectory detection algorithms that integrate the characteristics of different algorithms. Ge et al. proposed an anomaly trajectory detection algorithm from the aspects of direction and density. This algorithm integrates the idea of proximity on the basis of grids and reduces the influence of earlier arriving data on the evolving anomaly index through a decay function, and can effectively detect two types of anomaly trajectories based on direction and density, realizing the detection of evolving trajectory anomalies. However, the trajectory data of moving objects not only contains spatio-temporal features related to the object's position and time, but also contains multi-dimensional semantic features related to the object's motion characteristics and behaviors. The existing trajectory anomaly detection methods mainly consider the spatio-temporal features of trajectories and can detect anomaly trajectories under the constraints of spatio-temporal features, but ignore the influence of multi-dimensional semantic features of trajectories and cannot realize the detection of anomaly trajectories under the constraints of semantic features. Summary of the Invention
[0012] The purpose of the present invention is to solve the deficiencies existing in the prior art, and propose a semantic-aware method for detecting anomalies in trajectory data.
[0013] To achieve the above purpose, the present invention adopts the following technical solutions:
[0014] A semantic-aware method for detecting anomalies in trajectory data includes the following steps:
[0015] Step 1: Initialize the historical working data set T0, the abnormal point set X, the abnormal value score, and the adaptive window ω;
[0016] Step 2: Input the test trajectory. Based on the longitude, latitude, sampling time, and other information of the trajectory points, calculate the grid coordinates of the trajectory points and improve their semantic information;
[0017] Step 3: Add the trajectory points to the end of the adaptive window ω, filter the historical trajectory working set T according to the adaptive window i-1 , and generate a new working data set T i ;
[0018] Step 4: Calculate the support degree based on the newly obtained working data set T i and the original working data set T i-1 ;
[0019] Step 5: Calculate the current outlier based on the support degree and return to Step 3;
[0020] Step 6: Distinguish abnormal trajectory segments and abnormal trajectories. The abnormal segments in the trajectory can be accurately detected through the abnormal points in the outlier set. At the same time, the abnormal score of the trajectory can be used to sort the degree of trajectory abnormality to help identify abnormal trajectories.
[0021] Preferably, Step 1 includes the following steps:
[0022] Step 1.1: Initialize the outlier set X as empty;
[0023] Step 1.2: Initialize the abnormal score score as 0;
[0024] Step 1.3: Initialize the adaptive window ω as empty;
[0025] Step 1.4: Initialize the historical trajectory working set T0.
[0026] Preferably, Step 1.4 includes the following steps:
[0027] Step 1.4.1: Determine the grid size: Divide the target area evenly into grids. Calculate the grid side length based on the sampling time interval of the trajectory data and the average speed of the trajectory, and convert the grid side length into longitude difference and latitude difference. The calculation expression of the grid side length is as follows:
[0028] d = Δt × v'
[0029] where d represents the grid side length, Δt represents the sampling time interval, and v' represents the average speed of the trajectory;
[0030] Step 1.4.2: Map the trajectory to the grid: The trajectory consists of a series of GPS points with timestamps, where each GPS point contains longitude, latitude, sampling time, and speed information. Map the GPS points to the grid according to the longitude and latitude information of the GPS points, and convert the original trajectory point sequence into a grid sequence. The grid coordinate calculation expression is as follows:
[0031] x = (lon - mlon) / sizeX
[0032] y = (lat - mlat) / sizeY
[0033] Where x and y represent the grid abscissa and grid ordinate respectively, lon represents the longitude of the trajectory point, lat represents the latitude of the trajectory point, mlon represents the minimum longitude value in the historical trajectory working set, mlat represents the minimum latitude value in the historical working set, and sizeX and sizeY represent the longitude and latitude dimensions of a single grid;
[0034] Step 1.4.3: Improve the semantic information of the cells: Classify all POI information, and add the classified POI to the grid sequence according to the longitude and latitude information to improve the semantic information of the grid sequence;
[0035] Step 1.4.4: Generate an enhanced grid sequence: Since the GPS signal reception rate and the grid size do not match one by one, the mapped points in the grid sequence are not adjacent, resulting in the discontinuity of the grid sequence and leaving gaps. Supplement the gap area according to the line segment between two related cells to ensure that there are no gaps in the grid sequence, and calculate the relevant semantic information of the supplemented grid.
[0036] Preferably, the said Step 2 includes the following steps:
[0037] Step 2.1: Input the test trajectory and calculate the grid coordinates of the trajectory points: Calculate the grid coordinates of the trajectory points according to the longitude and latitude coordinates of the trajectory points, and map the trajectory points to the grid. The grid coordinate expression is as follows:
[0038] x = (lon - mlon) / sizeX
[0039] y = (lat - mlat) / sizeY
[0040] Where x and y represent the grid abscissa and grid ordinate respectively, lon represents the longitude of the trajectory point, lat represents the latitude of the trajectory point, mlon represents the minimum longitude value in the historical trajectory working set, mlat represents the minimum latitude value in the historical working set, and sizeX and sizeY represent the longitude and latitude dimensions of a single grid;
[0041] Step 2.2: Calculate the speed of the trajectory points: Calculate the speed of the trajectory points based on the longitude and latitude coordinates of the trajectory points and the sampling time interval. The expression for the speed of the trajectory points is as follows:
[0042]
[0043] where x and y represent the longitude and latitude coordinates of the trajectory points respectively, and Δt represents the sampling time interval;
[0044] Step 2.3: Improve the POI information: Add the classified POI information to the trajectory according to the longitude and latitude information to improve the POI information of each trajectory point.
[0045] Preferably, step 3 includes the following steps:
[0046] Step 3.1: Add the trajectory points of the test trajectory to the end of the adaptive window;
[0047] Step 3.2: Judge the length of the adaptive window. If it is 1, go to step 3.3; if it is greater than or equal to 2, go to step 3.4;
[0048] Step 3.3: Prune one of the historical working datasets;
[0049] Step 3.4: Prune the second of the historical working datasets.
[0050] Preferably, step 3.3 includes the following steps:
[0051] Step 3.3.1: Take out the cells from the grid sequence in the historical working dataset in turn;
[0052] Step 3.3.2: Calculate the similarity between the taken-out cell and the cell in the adaptive window using the edit distance, and divide all the information in the cell into spatial information, time information, and semantic information;
[0053] In the spatial information, due to the simplicity of the grid sequence augmentation method, the enhanced trajectory may not be completely accurate. Therefore, in the process of detecting the similarity of the spatial information of two cells, if the two cells are adjacent to each other, the edit distance between the two cells in the spatial information is considered to be 0, otherwise the edit distance is considered to be 1. The expression for the edit distance under spatial constraints is as follows:
[0054]
[0055] where g i represents the cell in the historical working dataset, and ω j represents the cell generated by the trajectory points in the adaptive window;
[0056] Strict equality is not required in terms of time and semantic information. If the difference between the two is within a certain range, the edit distance is considered to be 0; otherwise, it is 1. The edit distance expression under semantic conditions is as follows:
[0057]
[0058] where g ik represents the k-th dimensional semantics of the cell with a larger size in the historical work dataset, and ω ik represents the k-th dimensional semantics of the cell generated by the trajectory point in the adaptive window, and δ represents the threshold;
[0059] The edit distance expression between two cells is as follows:
[0060] EDR(g i , ω j ) = EDR(REST(g i , ω), REST(g i , ω)) + sub
[0061] where g i represents the cell in the historical work dataset, and ω represents the cell generated by the trajectory point in the adaptive window. If the two cells are similar under the constraint of the first-dimensional semantics, then sub = 0; otherwise, sub = 1;
[0062] Step 3.3.3: If the similarity between this cell and the cell in the adaptive window is greater than the threshold, then add this grid sequence to the new work dataset. The expression of the new work dataset is as follows:
[0063]
[0064] Preferably, the said Step 3.4 includes the following steps:
[0065] Step 3.4.1: Sequentially take out the cells from the grid sequences in the historical work dataset,
[0066] Step 3.4.2: Use the edit distance to calculate the similarity between the taken-out cell and the cell in the adaptive window. The edit distance expression between two cells is as follows:
[0067] EDR(g i , ω j ) = EDR(REST(g i , ω j ), REST(g i , ω j )) + sub
[0068] where g i represents the cell in the historical work dataset, and ωj Indicates the cells generated by the trajectory points in the adaptive window. If two cells are similar under the first-dimensional semantic condition constraint, then sub = 0; otherwise, sub = 1;
[0069] Step 3.4.3: If the similarity between this cell and the cells in the adaptive window is greater than the threshold, calculate the sequence number of this cell in the grid sequence;
[0070] Step 3.4.4: Calculate the sequence number of the cell similar to the penultimate cell in the adaptive window in the grid sequence;
[0071] Step 3.4.5: If two cells appear in the correct order in the grid sequence, add this grid sequence to the new working dataset, and its expression is as follows:
[0072]
[0073] where pos(t', g i ) represents the sequence number of grid g i appearing in trajectory t'.
[0074] Preferably, step 4 includes the following steps:
[0075] Step 4.1: Calculate the sizes of the historical working dataset and the new working dataset;
[0076] Step 4.2: Calculate the support degree. The support degree expression is as follows:
[0077] support(i) = count(hasPath(T i-1 , ω)) / count(T i-1 )
[0078] where count(T) represents the number of trajectories in the working dataset.
[0079] Preferably, step 5 includes the following steps:
[0080] Step 5.1: Calculate the current outlier value. Since the trajectory is ongoing, an outlier score will be retained to provide an alarm during the detection process and at the end of the detection. After the trajectory detection is completed, it will be sorted. Among them, the smaller the support degree and the longer the abnormal distance, the higher the ranking of the trajectory. Therefore, this score is calculated according to the length of the abnormal sub-segment and the density in each abnormal sub-segment, and its expression is as follows:
[0081]
[0082] where λ is the heat parameter used to smooth the breakpoints caused by the threshold selection, dist(pi , p i-1 ) represents the distance between two adjacent trajectory points in the test trajectory;
[0083] Step 5.2: Determine whether the support is greater than the threshold. If it is greater than the threshold, go to Step 3; if it is less than the threshold, go to Step 5.3;
[0084] Step 5.3: Reset the historical working data set and the adaptive window.
[0085] Preferably, Step 5.3 includes the following steps:
[0086] Step 5.3.1: Add the detection point to the abnormal point set and calculate the abnormal value;
[0087] Step 5.3.2: Reset the historical trajectory working set to the original historical trajectory working set;
[0088] Step 5.3.2: Reset the adaptive window and only keep the last added grid;
[0089] Step 5.3.4: Go to Step 3.
[0090] Compared with the prior art, the beneficial effects of the present invention are:
[0091] 1. Aiming at the problem that it is difficult to determine the grid size in the grid-based abnormal trajectory detection method, the grid size is determined by the sampling interval of trajectory points and the average movement speed of trajectory points, which can effectively avoid problems such as insufficient detection accuracy and excessive calculation amount caused by too large or too small grids.
[0092] 2. Aiming at the problem that the existing trajectory anomaly detection methods mainly consider spatio-temporal features and ignore the multi-dimensional semantic features of trajectories, and cannot achieve abnormal trajectory detection under multi-dimensional semantic feature constraints. The present invention uses the edit distance to calculate the similarity between cells, and can achieve abnormal trajectory detection under multi-dimensional semantic constraints. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 It is a flow chart of the abnormal trajectory detection method considering semantics of the present invention;
[0094] Figure 2 It is the cell semantic vector space in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0095] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0096] Embodiment 1
[0097] A semantic-aware abnormal detection method for trajectory data, comprising the following steps:
[0098] Step 1: Initialize the historical work dataset T0, the abnormal point set X, the abnormal score, and the adaptive window ω;
[0099] The said Step 1 comprises the following steps:
[0100] Step 1.1: Initialize the abnormal point set X as empty;
[0101] Step 1.2: Initialize the abnormal score as 0;
[0102] Step 1.3: Initialize the adaptive window ω as empty;
[0103] Step 1.4: Initialize the historical trajectory work set T0.
[0104] The said Step 1.4 comprises the following steps:
[0105] Step 1.4.1: Determine the grid size: evenly divide the target area into grids, calculate the grid side length according to the sampling time interval of the trajectory data and the average speed of the trajectory, and convert the grid side length into longitude difference and latitude difference. The calculation expression of the grid side length is as follows:
[0106] d = Δt × v'
[0107] where d represents the grid side length, Δt represents the sampling time interval, and v' represents the average speed of the trajectory;
[0108] Step 1.4.2: Map the trajectory to the grid: The trajectory is composed of a series of GPS points with timestamps, where each GPS point contains longitude and latitude, sampling time, and speed information. Map the GPS points to the grid according to the longitude and latitude information of the GPS points, and convert the original trajectory point sequence into a grid sequence. The calculation expressions of the grid coordinates are as follows:
[0109] x = (lon - mlon) / sizeX
[0110] y = (lat - mlat) / sizeY
[0111] where x and y respectively represent the grid abscissa and grid ordinate, lon represents the longitude of the trajectory point, lat represents the latitude of the trajectory point, mlon represents the minimum longitude value in the historical trajectory work set, mlat represents the minimum latitude value in the historical work set, and sizeX and sizeY represent the longitude and latitude sizes of a single grid;
[0112] Step 1.4.3: Improve the semantic information of cells: Classify all POI information, add the classified POIs to the grid sequence according to the longitude and latitude information, and improve the semantic information of the grid sequence;
[0113] Step 1.4.4: Generate an enhanced grid sequence: Since the GPS signal reception rate and the grid size do not match one by one, the mapping points in the grid sequence are not adjacent, resulting in the discontinuity of the grid sequence and leaving gaps. Supplement the gap area according to the line segment between two related cells to ensure that there are no gaps in the grid sequence, and calculate the relevant semantic information of the supplemented grid.
[0114] Step 2: Input the test trajectory, and calculate the grid coordinates of the trajectory points and improve their semantic information according to the longitude, latitude, sampling time and other information of the trajectory points;
[0115] The said Step 2 includes the following steps:
[0116] Step 2.1: Input the test trajectory and calculate the grid coordinates of the trajectory points: Calculate the grid coordinates of the trajectory points according to the longitude and latitude coordinates of the trajectory points, and map the trajectory points to the grid. The grid coordinate expression is as follows:
[0117] x = (lon - mlon) / sizeX
[0118] y = (lat - mlat) / sizeY
[0119] where x and y respectively represent the grid abscissa and grid ordinate, lon represents the longitude of the trajectory point, lat represents the latitude of the trajectory point, mlon represents the minimum longitude value in the historical trajectory working set, mlat represents the minimum latitude value in the historical working set, and sizeX and sizeY represent the longitude and latitude sizes of a single grid;
[0120] Step 2.2: Calculate the speed of the trajectory points: Calculate the speed of the trajectory points according to the longitude and latitude coordinates of the trajectory points and the sampling time interval. The trajectory point speed expression is as follows:
[0121]
[0122] where x and y respectively represent the longitude and latitude coordinates of the trajectory point, and Δt represents the sampling time interval;
[0123] Step 2.3: Improve the POI information: Add the classified POI information to the trajectory according to the longitude and latitude information, and improve the POI information of each trajectory point.
[0124] Step 3: Add the trajectory points to the end of the adaptive window ω, filter the historical trajectory working set T according to the adaptive window i-1 , and generate a new working data set T i ;
[0125] Step 3 includes the following steps:
[0126] Step 3.1: Add the trajectory points of the test trajectory to the end of the adaptive window;
[0127] Step 3.2: Judge the length of the adaptive window. If it is 1, go to Step 3.3; if it is greater than or equal to 2, go to Step 3.4;
[0128] Step 3.3: Prune one of the historical working data sets;
[0129] Step 3.4: Prune the second of the historical working data sets.
[0130] The said Step 3.3 includes the following steps:
[0131] Step 3.3.1: Take out the cells from the grid sequence in the historical working data set in turn;
[0132] Step 3.3.2: Calculate the similarity between the taken-out cell and the cell in the adaptive window using the edit distance, and divide all the information in the cell into spatial information, time information, and semantic information;
[0133] In the spatial information, due to the simplicity of the grid sequence augmentation method, the enhanced trajectory may not be completely accurate. Therefore, in the process of detecting the similarity of the spatial information of two cells, if the two cells are adjacent cells to each other, it is considered that the edit distance between the two cells in the spatial information is 0, otherwise the edit distance is 1. The edit distance expression under the spatial constraint is as follows:
[0134]
[0135] where g i represents the cell in the historical working data set, and ω j represents the cell generated by the trajectory point in the adaptive window;
[0136] In the time and semantic information, strict equality is not required. If the difference between the two is within a certain range, the edit distance is considered to be 0, otherwise it is 1. The edit distance expression under the semantic condition constraint is as follows:
[0137]
[0138] where g ik represents the k-th dimensional semantics of the larger cell in the historical working data set, and ω ik represents the k-th dimensional semantics of the cell generated by the trajectory point in the adaptive window, and δ represents the threshold;
[0139] The edit distance expression between two cells is as follows:
[0140] EDR(g i , ω j )) = EDR(REST(g i , ω), REST(g i , ω)) + sub
[0141] where g i represents a cell in the historical working dataset, and ω represents a cell generated by a trajectory point in the adaptive window. If the two cells are similar under the first - dimension semantic condition constraint, then sub = 0; otherwise, sub = 1;
[0142] Step 3.3.3: If the similarity between this cell and the cell in the adaptive window is greater than the threshold, then add this grid sequence to the new working dataset. The expression of the new working dataset is as follows:
[0143]
[0144] The said Step 3.4 includes the following steps:
[0145] Step 3.4.1: Take out cells from the grid sequences in the historical working dataset in turn,
[0146] Step 3.4.2: Calculate the similarity between the taken - out cell and the cell in the adaptive window using the edit distance. The expression of the edit distance between two cells is as follows:
[0147] EDR(g i , ω j ) = EDR(REST(g i , ω j ), REST(g i , ω j )) + sub
[0148] where g i represents a cell in the historical working dataset, and ω j represents a cell generated by a trajectory point in the adaptive window. If the two cells are similar under the first - dimension semantic condition constraint, then sub = 0; otherwise, sub = 1;
[0149] Step 3.4.3: If the similarity between this cell and the cell in the adaptive window is greater than the threshold, then calculate the serial number of this cell in the grid sequence;
[0150] Step 3.4.4: Calculate the serial number of the cell similar to the penultimate cell in the adaptive window in the grid sequence;
[0151] Step 3.4.5: If two cells appear in the correct order in the grid sequence, add the grid sequence to the new working dataset, and its expression is as follows:
[0152]
[0153] where pos(t', g i ) represents the serial number of grid g i appearing in the trajectory t'.
[0154] Step 4: Calculate the support degree according to the new working dataset T i calculated previously and the original working dataset T i-1 ; Step 4 includes the following steps:
[0155] Step 4.1: Calculate the sizes of the historical working dataset and the new working dataset;
[0156] Step 4.2: Calculate the support degree. The support degree expression is as follows:
[0157] support(i) = count(hasPath(T i-1 , ω)) / count(T i-1 )
[0158] where count(T) represents the number of trajectories in the working dataset.
[0159] Step 5: Calculate the current outlier according to the support degree, and return to Step 3;
[0160] Step 5 includes the following steps:
[0161] Step 5.1: Calculate the current outlier. Since the trajectory is ongoing, an outlier score will be retained to provide an alarm during the detection process and provide an alarm when the detection is completed, and it will be sorted after the trajectory detection is completed. Among them, the smaller the support degree and the longer the abnormal distance, the higher the ranking of the trajectory. Therefore, calculate this score according to the length of the abnormal sub-segment and the density in each abnormal sub-segment, and its expression is as follows:
[0162]
[0163] where λ is the heat parameter used to smooth the breakpoints caused by the threshold selection, and dist(p i , p i-1 ) represents the distance between two adjacent trajectory points in the test trajectory;
[0164] Step 5.2: Determine whether the support degree is greater than the threshold. If it is greater than the threshold, go to Step 3; if it is less than the threshold, go to Step 5.3;
[0165] Step 5.3: Reset the historical working data set and the adaptive window.
[0166] The said step 5.3 includes the following steps:
[0167] Step 5.3.1: Add the detection point to the outlier set and calculate the outlier value;
[0168] Step 5.3.2: Reset the historical trajectory working set to the original historical trajectory working set;
[0169] Step 5.3.3: Reset the adaptive window, only keeping the last added grid;
[0170] Step 5.3.4: Go to step 3.
[0171] Step 6: Distinguish the abnormal trajectory segments and abnormal trajectories. The abnormal segments in the trajectory can be accurately detected through the outliers in the outlier set, and at the same time, the abnormal degree of the trajectory can be ranked through the abnormal score of the trajectory to help identify the abnormal trajectories.
[0172] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A semantic-aware trajectory data anomaly detection method, characterized in that: Including the following steps: Step 1: Initialize the historical work dataset , the outlier set , the outlier value and the adaptive window ; Step 2: Input the test trajectory. Based on the longitude and latitude of the trajectory points and the sampling time information, calculate the grid coordinates of the trajectory points and improve their semantic information; The said Step 2 includes the following steps: Step 2.1: Input the test trajectory and calculate the grid coordinates of the trajectory points: Calculate the grid coordinates of the trajectory points according to the longitude and latitude coordinates of the trajectory points, and map the trajectory points to the grid. The grid coordinate expression is as follows: ; Among them , respectively represent the grid abscissa and grid ordinate, represents the longitude of the trajectory point, represents the latitude of the trajectory point, represents the minimum longitude value in the historical trajectory working set, represents the minimum latitude value in the historical working set, , represents the longitude and latitude size of a single grid; Step 2.2: Calculate the speed of the trajectory points: Calculate the speed of the trajectory points according to the longitude and latitude coordinates of the trajectory points and the sampling time interval. The trajectory point speed expression is as follows: ; wherein , respectively represent the longitude and latitude coordinates of the trajectory points, represents the sampling time interval; Step 2.3: Improve the POI information: Add the classified POI information to the trajectory according to the longitude and latitude information to improve the POI information of each trajectory point; Step 3: Add the trajectory points to the adaptive window At the end, filter the historical trajectory working set according to the adaptive window and generate a new working data set ; Step 4: Calculate the support based on the newly obtained working data set calculated previously and the original working data set Step 5: Calculate the current outlier according to the support degree and return to Step 3; The said Step 5 includes the following steps: Step 5.1: Calculate the current outlier. Since the trajectory is ongoing, an outlier score will be retained to provide an alarm during the detection process and also provide an alarm when the detection is completed, and sort it after the trajectory detection is completed. Among them, the smaller the support degree and the longer the abnormal distance, the higher the ranking of the trajectory. Therefore, calculate this outlier score according to the length of the abnormal sub-segment and the density in each abnormal sub-segment. The expression is as follows: ; wherein is the heat parameter, used to smooth the breakpoints caused by threshold selection, represents the distance between two adjacent trajectory points in the test trajectory; Step 5.2: Judge whether the support degree is greater than the threshold. If it is greater than the threshold, go to Step 3. If it is less than the threshold, go to Step 5.3; Step 5.3: Reset the historical working data set and the adaptive window; Step 6: Distinguish the abnormal trajectory segments and abnormal trajectories. The abnormal segments in the trajectory can be accurately detected through the abnormal points in the outlier set. At the same time, the abnormal degree of the trajectory can be sorted through the outlier score of the trajectory to help identify the abnormal trajectories.
2. The semantic-aware trajectory data anomaly detection method according to claim 1, wherein: The said Step 1 includes the following steps: Step 1.1: Initialize the set of outlier points to be empty; Step 1.2: Initialize the anomaly score to 0; Step 1.3: Initialize the adaptive window to be empty; Step 1.4: Initialize the historical trajectory working set .
3. The method for detecting abnormal trajectory data considering semantics according to claim 2, characterized in that: The said Step 1.4 includes the following steps: Step 1.4.1: Determine the grid size: Divide the target area evenly into grids, calculate the grid side length according to the sampling time interval of the trajectory data and the average speed of the trajectory, and convert the grid side length into the longitude difference and latitude difference. The grid side length calculation expression is as follows: ; Among them represents the grid side length, represents the sampling time interval, represents the average track speed; Step 1.4.2: Map the trajectory to the grid: The trajectory is composed of a series of GPS points with timestamps, where each GPS point contains longitude and latitude, sampling time, and speed information. Map the GPS points to the grid according to the longitude and latitude information of the GPS points, and convert the original trajectory point sequence into a grid sequence. The grid coordinate calculation expression is as follows: ; Among them , respectively represent the abscissa and ordinate of the grid, represents the longitude of the trajectory point, represents the latitude of the trajectory point, represents the minimum longitude value in the historical trajectory working set, represents the minimum latitude value in the historical working set, , represents the longitude and latitude size of a single grid; Step 1.4.3: Improve the semantic information of the cells: Classify all the POI information, and add the classified POI to the grid sequence according to the longitude and latitude information to improve the semantic information of the grid sequence; Step 1.4.4: Generate an enhanced grid sequence: Since the GPS signal reception rate and the grid size do not match one by one, the mapped points in the grid sequence are not adjacent, resulting in the discontinuity of the grid sequence and leaving gaps. Supplement the gap area according to the line segment between two related cells to ensure that there are no gaps in the grid sequence, and at the same time calculate the relevant semantic information of the supplemented grid.
4. The semantic-aware trajectory data anomaly detection method according to claim 1, wherein: The said Step 3 includes the following steps: Step 3.1: Add the track points of the test track to the end of the adaptive window; Step 3.2: Judge the length of the adaptive window. If it is 1, go to Step 3.3; if it is greater than or equal to 2, go to Step 3.4; Step 3.3: Prune one of the historical working data sets; Step 3.4: Prune the second of the historical working data sets.
5. A method for detecting abnormal trajectory data considering semantics according to claim 4, characterized in that: The said Step 3.3 includes the following steps: Step 3.3.1: Take out cells from the grid sequence in the historical working data set in turn; Step 3.3.2: Calculate the similarity between the taken-out cell and the cell in the adaptive window using the edit distance, and divide all the information in the cell into spatial information, time information, and semantic information; In the spatial information, due to the simplicity of the grid sequence augmentation method, the enhanced track may not be completely accurate. Therefore, in the process of detecting the similarity of the spatial information of two cells, if the two cells are adjacent cells to each other, it is considered that the edit distance of the two cells in the spatial information is 0, otherwise the edit distance is 1. The edit distance expression under the spatial constraint is as follows: ; Among them represents a cell in the historical work dataset, represents a cell generated by trajectory points in the adaptive window; In the time and semantic information, strict equality is not required. If the difference between the two is within a certain range, the edit distance is considered 0, otherwise it is 1. The edit distance expression under the semantic condition constraint is as follows: ; wherein represents the k-th dimensional semantics of the cell in the historical working dataset with a large size, represents the k-th dimensional semantics of the cell generated by the trajectory points in the adaptive window, represents a threshold value; The edit distance expression between two cells is as follows: ; Among them represents the cell in the historical work dataset represents the cell generated by the trajectory points in the adaptive window. If two cells are similar under the first-dimensional semantic condition constraint, then , otherwise ; Step 3.3.3: If the similarity between the cell and the cell in the adaptive window is greater than the threshold, add the grid sequence to the new working data set. The expression of the new working data set is as follows: 。 6. The method for detecting abnormal trajectory data considering semantics according to claim 5, characterized in that: The said Step 3.4 includes the following steps: Step 3.4.1: Take out cells from the grid sequence in the historical working data set in turn, Step 3.4.2: Calculate the similarity between the taken-out cell and the cell in the adaptive window using the edit distance. The edit distance expression between two cells is as follows: ; Among them represents the cell in the historical work dataset represents the cell generated by the trajectory points in the adaptive window. If two cells are similar under the constraint of the first-dimensional semantic condition, then , otherwise ; Step 3.4.3: If the similarity between the cell and the cell in the adaptive window is greater than the threshold, calculate the serial number of the cell in the grid sequence; Step 3.4.4: Calculate the serial number of the cell similar to the penultimate cell in the adaptive window in the grid sequence; Step 3.4.5: If the two cells appear in the correct order in the grid sequence, add the grid sequence to the new working data set. Its expression is as follows: ; Among them represents the grid in the trajectory the serial number that appears 7. A semantic-aware trajectory data anomaly detection method according to claim 1, characterized in that: The said Step 4 includes the following steps: Step 4.1: Calculate the sizes of the historical working data set and the new working data set; Step 4.2: Calculate the support degree. The support degree expression is as follows: ; Among them represents the number of trajectories in the working dataset.
8. The semantic-aware trajectory data anomaly detection method according to claim 1, characterized in that: The said Step 5.3 includes the following steps: Step 5.3.1: Add the detection point to the abnormal point set and calculate the abnormal value; Step 5.3.2: Reset the historical track working set to the original historical track working set; Step 5.3.2: Reset the adaptive window, only keep the last added grid; Step 5.3.4: Go to Step 3.
Citation Information
Patent Citations
Cluster communication terminal track real time anomaly detection method and system based on hybrid grid hierarchical clustering
CN105825242A
Autonomous nap-of-the-earth (ANOE) flight path planning for manned and unmanned rotorcraft
US20160210863A1