Adaptive Trajectory Inflection Point Extraction and Compression Method and System Based on Partitioning

Through the adaptive trajectory inflection point extraction and compression method, the problems of trajectory data scale explosion and threshold selection are solved, and efficient trajectory compression and accurate inflection point recognition are achieved.

CN116483814BActive Publication Date: 2025-05-30FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310449353.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-05-30
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

In the prior art, the high sampling frequency of trajectory data leads to an explosion in data scale, affecting the efficiency of storage, management and query analysis, and most compression algorithms require manual selection of thresholds and lack adaptability.

Method used

A method of adaptive trajectory inflection point extraction and compression based on division is proposed. By acquiring trajectory data, calculating direction vectors, using Kmeans clustering and cosine similarity to trajectory division and merging, combining the improved contour coefficient to judge the boundary movement of sub-trajectory, extracting the starting point, inflection point and end point of the trajectory to generate the compressed trajectory.

Benefits of technology

High-precision adaptive trajectory inflection point extraction is realized, reducing the storage space of trajectory data, improving data analysis efficiency, and avoiding the defect of artificially selecting thresholds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116483814B_ABST
    Figure CN116483814B_ABST
Patent Text Reader

Abstract

The object of the present invention is to provide a method and system for extracting and compressing adaptive trajectory inflection points based on partitioning, by constructing a basic trajectory data set; calculating the driving direction vectors of two adjacent trajectory points in sequence to obtain a direction vector sequence of the trajectory. Clustering the direction vector sequence based on the Kmeans algorithm, roughly dividing the trajectory into sub-trajectories to obtain a sub-trajectory sequence; searching for sub-trajectories containing only two trajectory points in the sub-trajectory sequence as sub-trajectories to be merged, calculating the merging threshold of each sub-trajectory to be merged based on cosine similarity, and merging the sub-trajectories that need to be merged; judging the conditions for the boundary movement of the sub-trajectory based on the improved silhouette coefficient, moving the boundary of the sub-trajectory that needs to be moved, and adjacent sub-trajectories share a trajectory point, extracting such trajectory points as the inflection points of the trajectory; extracting the starting point, inflection point and ending point of the trajectory sequence, and the trajectory sequence is the compressed trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of trajectory storage management, and particularly relates to a method and system for extracting and compressing adaptive trajectory inflection points based on partitioning. Background Art

[0002] With the rapid development and widespread application of positioning technology and mobile communication technology, the acquisition of trajectory data has become increasingly easy, and important research value has emerged in different fields. However, a high sampling frequency will generate a large number of trajectory records, resulting in an explosive growth in the scale of trajectory data, which will seriously affect the efficiency of data storage, management, query and analysis, bring great pressure to the server, and pose huge challenges to further data mining work. To reduce the storage space of trajectory data and simplify trajectory data analysis, the compressed storage of trajectory data has become a way and research hotspot to accelerate trajectory pattern mining. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method and system for extracting and compressing adaptive trajectory inflection points based on partitioning, which can achieve high-precision adaptive trajectory inflection point extraction and efficient trajectory compression.

[0004] Aiming at the problem that most current compression algorithms require artificial selection of thresholds, the present invention starts from the global direction feature and local direction feature of the trajectory, and proposes a completely adaptive trajectory inflection point extraction and compression algorithm to achieve high-precision adaptive trajectory inflection point extraction.

[0005] The main process of the method includes: obtaining trajectory data, removing duplicate records and abnormal data and complementing missing data to construct a basic trajectory data set; calculating the driving direction vectors of two adjacent trajectory points in sequence to obtain a sequence of direction vectors of the trajectory. Based on the Kmeans algorithm, clustering is performed on the sequence of direction vectors to roughly divide the trajectory into sub-trajectories and obtain a sequence of sub-trajectories; searching for sub-trajectories containing only two trajectory points in the sequence of sub-trajectories as sub-trajectories to be merged, calculating the merging threshold of each sub-trajectory to be merged based on cosine similarity, and merging the sub-trajectories that need to be merged to obtain a merged sequence of sub-trajectories; based on the improved silhouette coefficient, judging the conditions for the movement of the sub-trajectory boundary, moving the boundary of the sub-trajectory that needs to be moved to obtain an updated sequence of sub-trajectories, where adjacent sub-trajectories share a trajectory point, and extracting such trajectory points as the inflection points of the trajectory; extracting the starting point, inflection points and ending point of the trajectory sequence, and the trajectory sequence is the compressed trajectory. The present invention has the advantages of adaptive and high-precision inflection point recognition, has high reference value in the scenario of trajectory compression, and can also be applied to the work of trajectory partitioning based on direction.

[0006] The technical solution specifically adopted by the present invention to solve its technical problems is:

[0007] An adaptive trajectory inflection point extraction and compression method based on partitioning, comprising the following steps:

[0008] Step S1: Obtain trajectory data, and after preprocessing, construct a basic trajectory data set;

[0009] Step S2: Calculate the driving direction vectors of two adjacent trajectory points in sequence to obtain a sequence of direction vectors of the trajectory; perform clustering on the sequence of direction vectors based on the Kmeans algorithm, roughly divide the trajectory into sub-trajectories, and obtain a sequence of sub-trajectories;

[0010] Step S3: Search for sub-trajectories containing only two trajectory points in the sequence of sub-trajectories as sub-trajectories to be merged, calculate the merging threshold of each sub-trajectory to be merged based on cosine similarity, and merge the sub-trajectories that need to be merged to obtain a merged sequence of sub-trajectories;

[0011] Step S4: Based on the improved silhouette coefficient, judge the conditions for the movement of the sub-trajectory boundary, move the boundary of the sub-trajectory that needs to be moved, and obtain an updated sequence of sub-trajectories. Adjacent sub-trajectories share a trajectory point, which serves as the inflection point of the trajectory;

[0012] Step S5: Extract the starting point, inflection points, and ending point of the trajectory sequence, and the generated trajectory sequence is the compressed trajectory.

[0013] Further, the preprocessing in step S1 includes: removing duplicate records and abnormal data and filling in missing data.

[0014] Further, step S2 specifically includes the following steps:

[0015] Step S21: Calculate the direction of each pair of consecutive trajectory points in sequence, and map a trajectory to a sequence of direction vectors; the calculation formula is as follows:

[0016]

[0017] In the formula, p describes the trajectory points in the trajectory, lon and lat are the longitude and latitude of point p respectively; a sequence of direction vectors is described as

[0018] Step S22: Divide the driving direction of the trajectory into eight core directions, with the directions spaced 45° apart from each other. The eight core directions divide the driving direction into eight intervals; record the interval to which each direction vector in the sequence of direction vectors belongs, and count the total number of intervals to which they belong as the input parameter k required by the Kmeans clustering algorithm. The value range of k is [2, 8]. When all direction vectors are in the same interval, k is taken as 2;

[0019] Step S23: Cluster the trajectory direction vectors and use the cosine distance for distance measurement; according to the clustering results, assign the same label to the direction vectors belonging to the same class, and divide the trajectory at the label change to achieve a rough division of the trajectory, obtaining a subsequence of trajectories.

[0020] Further, step S3 specifically includes the following steps:

[0021] Step S31: Calculate a merging threshold for each sub-trajectory to be merged; the process description is as follows: Calculate the cosine distance between the direction vector corresponding to the sub-trajectory and other direction vectors belonging to the same class, and take the maximum value as the merging threshold; the calculation formula is as follows:

[0022]

[0023] In the formula, x is described as the unique direction vector of the sub-trajectory to be merged, and x i is other direction vectors belonging to the same cluster as x, and the similarity measurement between vectors uses the cosine distance;

[0024] Step S32: Calculate the average value of the cosine distances between the direction of each sub-trajectory to be merged and all the direction vectors of the adjacent sub-trajectories. If it is greater than the merging threshold, merge the sub-trajectory with the adjacent sub-trajectory with a more similar direction to obtain the merged subsequence of sub-trajectories.

[0025] Further, step S4 specifically includes the following steps:

[0026] Considering the temporal characteristics of the trajectory data, calculate the silhouette coefficient at the boundary trajectory segments of the sub-trajectories, and determine whether to move the boundary of the sub-trajectory and the direction of the boundary movement according to the size of the silhouette coefficient; the calculation formula is as follows:

[0027]

[0028] Where:

[0029]

[0030] In the formula, the silhouette coefficient has a value between [-1, 1], the closer it is to 1, the relatively better the cohesion and separation degree of the direction vector are; conversely, the closer it is to -1, the relatively worse the cohesion and separation degree of the direction vector are; is to calculate the cohesion degree of the direction of the trajectory segment, that is, it is described as the average value of the dissimilarity degree between the current direction vector and other direction vectors within the same sub-trajectory; It is to calculate the external separation degree of the current direction vector, which is described as the current direction vector and the average value of the dissimilarity degree with all direction vectors in its directly adjacent sub-trajectory;

[0031] During the process of boundary movement, if the current sub-trajectory only contains two trajectory points after the movement, the judgment of the merging condition is performed on this sub-trajectory, and the sub-trajectories that need to be merged are merged.

[0032] The sub-trajectory sequence after boundary movement is obtained. Adjacent sub-trajectories share a trajectory point, and such trajectory points are extracted as the inflection points of the trajectory.

[0033] And, an adaptive trajectory inflection point extraction and compression system based on partitioning, characterized by including:

[0034] A trajectory rough partitioning module, which is used to obtain a sequence of trajectory direction vectors by calculating the driving direction of the trajectory, and use the Kmeans clustering algorithm to cluster the vectors, so as to explore the direction distribution and changes in the trajectory from a global perspective, so as to obtain a preliminary trajectory partitioning result;

[0035] A sub-trajectory merging module, which is used to adjust the result of the trajectory rough partitioning from a local perspective, and appropriately merge some sub-trajectories based on cosine similarity, so as to effectively reduce potential incorrect trajectory partitioning results and improve the accuracy of trajectory partitioning;

[0036] A trajectory fine partitioning module, which is used to adjust the result of the trajectory partitioning from a local perspective again, improve the silhouette coefficient, and evaluate the result of the trajectory partitioning from two aspects of the homogeneity within the sub-trajectory and the heterogeneity between the sub-trajectories, so as to further improve the accuracy of the trajectory partitioning.

[0037] Compared with the prior art, the present invention and its preferred solutions integrate the idea of trajectory direction feature clustering and the unsupervised sequence partitioning method: on the one hand, by clustering the trajectory directions, potential trajectory inflection points are mined from a global perspective; on the other hand, by referring to the sequence partitioning method, partitioning is performed at possible trajectory inflection points, and the partitioning results are appropriately merged and adjusted, so as to further improve the accuracy of trajectory inflection point recognition from a local perspective. The algorithm has the advantages of adaptive and high-precision inflection point recognition, and has high reference value in the scenario of trajectory compression. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0039] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention;

[0040] Figure 2It is a comparison chart of the running time consumption of the method obtained in the embodiment of the present invention and other methods;

[0041] Figure 3 It is a comparison chart of the compression ratios of the method obtained in the embodiment of the present invention and other methods under the same direction error level;

[0042] Figure 4 It is a comparison chart of the average direction errors of the method obtained in the embodiment of the present invention and other methods under the same compression ratio level;

[0043] Figure 5 It is a comparison chart of the visual results of representative trajectory inflection point extraction of the method obtained in the embodiment of the present invention and other methods;

[0044] In the figure: a) Comparison of the recognition results of representative trajectory inflection points in the short trajectory dataset; b) Comparison of the recognition results of representative trajectory inflection points in the medium and long trajectory datasets; c) Comparison of the recognition results of representative trajectory inflection points in the long trajectory dataset; d) Comparison of the recognition results of representative trajectory inflection points in the ultra-long trajectory dataset. Detailed implementation manners

[0045] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0046] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further explanations for the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0047] It should be noted that the terms used here are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used here, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0048] As Figure 1 shown, the present invention provides a trajectory inflection point extraction and compression algorithm based on partitioning, including the following steps:

[0049] Step S1: Obtain trajectory data, preprocess it, and construct a basic trajectory dataset;

[0050] Step S2: Calculate the direction of the trajectory to obtain a sequence of direction vectors of the trajectory, and divide the trajectory into several sub-trajectories based on Kmeans clustering to obtain a sequence of sub-trajectories;

[0051] Step S3: Search for sub-trajectories that contain only two trajectory points in the sub-trajectory sequence as the sub-trajectories to be merged. Calculate the merging threshold for each sub-trajectory to be merged based on cosine similarity, and merge the sub-trajectories that need to be merged to obtain the merged sub-trajectory sequence;

[0052] Step S4: Based on the improved silhouette coefficient, judge the conditions for the movement of the sub-trajectory boundary, move the boundary of the sub-trajectory that needs to be moved to obtain the updated sub-trajectory sequence. Adjacent sub-trajectories share a trajectory point, and extract such trajectory points as the inflection points of the trajectory;

[0053] Step S5: Extract the starting point, inflection point and ending point of the trajectory sequence, and the trajectory sequence is the compressed trajectory.

[0054] Further, step S2 is specifically as follows:

[0055] Step S21: Calculate the direction of each pair of consecutive trajectory points in turn, and map a trajectory into a sequence of direction vectors; the calculation formula is as follows:

[0056]

[0057] In the formula, p describes the trajectory points in the trajectory, and lon and lat are the longitude and latitude of point p respectively. A sequence of direction vectors can be described as

[0058] Step S22: Divide the driving direction of the trajectory into eight core directions, with an interval of 45° between each direction. The eight core directions exactly divide the driving direction into eight intervals; record the interval to which each direction vector in the sequence of direction vectors belongs, and count the total number of intervals as the input parameter k required by the Kmeans clustering algorithm. The value range of k is [2, 8], and k is taken as 2 when all direction vectors are in the same interval;

[0059] Step S23: Cluster the trajectory direction vectors, and use the cosine distance instead of the traditional Euclidean distance for distance measurement. According to the clustering results, the direction vectors belonging to the same class will be assigned the same label, and the trajectory is roughly divided at the label change to obtain the sub-trajectory sequence.

[0060] Further, step S3 is specifically as follows:

[0061] Step S31: Calculate a merging threshold for each sub-trajectory to be merged. The process is described as calculating the cosine distance between this direction vector and other direction vectors belonging to the same class, and taking the maximum value as the merging threshold; the calculation formula is as follows:

[0062]

[0063] Wherein, x is described as the unique direction vector of the sub-trajectory to be merged, and x i is other direction vectors belonging to the same cluster as x, and the cosine distance is used as the similarity measure between vectors.

[0064] Step S32: Calculate the average value of the cosine distances between the direction of each sub-trajectory to be merged and all direction vectors of adjacent sub-trajectories. If it is greater than the merging threshold, merge the sub-trajectory with the adjacent sub-trajectory with a more similar direction to obtain the merged sub-trajectory sequence.

[0065] Furthermore, step S4 is specifically as follows:

[0066] Step S41: Use the silhouette coefficient to evaluate the effect of trajectory partitioning from two aspects: the homogeneity within the sub-trajectory and the heterogeneity between sub-trajectories. Considering the characteristics of time series data, calculate the silhouette coefficient only at the boundary trajectory segments of the sub-trajectory, and determine whether to move the boundary of the sub-trajectory and the direction of the boundary movement according to the size of the silhouette coefficient; the calculation formula is as follows:

[0067]

[0068] Wherein:

[0069]

[0070] In the formula, the silhouette coefficient has a value between [-1, 1], the closer it is to 1, the relatively better the cohesion and separation of the direction vector are; conversely, the closer it is to -1, the relatively worse the cohesion and separation of the direction vector are. is to calculate the cohesion (homogeneity) of the direction of the trajectory segment, that is, it is described as the average value of the dissimilarity degree between the current direction vector and other direction vectors within the same sub-trajectory; is to calculate the external separation (heterogeneity) of the current direction vector, which is described as the average value of the dissimilarity degree between the current direction vector and all direction vectors within its directly adjacent sub-trajectory.

[0071] Step S42: During the process of boundary movement, if it happens that the current sub-trajectory only contains two trajectory points after the movement, then judge the merging conditions for this sub-trajectory and merge the sub-trajectories that need to be merged.

[0072] Step S43: Obtain the sub-trajectory sequence after boundary movement. Adjacent sub-trajectories share a trajectory point, and extract such trajectory points as the inflection points of the trajectory.

[0073] Example 1:

[0074] This example uses the GPS trajectories collected in the Geolife project by 182 users over a period of more than five years (from April 2007 to August 2012). The GPS trajectories in this dataset are represented by a series of sampling time points, each point containing latitude, longitude, and altitude information. The trajectories have a high sampling rate, and 91.5% of the trajectories have a high sampling density. The specific implementation is as follows:

[0075] Step S1: Obtain the trajectory attribute data of moving objects in the Geolife project, directly eliminate duplicate records and abnormal data, and complete the missing data to obtain a basic experimental trajectory dataset; divide the trajectory dataset into different levels according to the number of points in the trajectory, and construct a hierarchical experimental trajectory dataset.

[0076] Table 1

[0077] Track dataset level Number of tracks (pieces) Number of track points (pieces) Short track 2102 10~300 Medium and long track 1430 300~1000 Long track 409 1000~2000 Ultra-long track 247 >2000

[0078] As shown in Table 1, the detailed information of the trajectory dataset is divided into four levels according to the number of points in the trajectory, namely short trajectories (10 - 300 points), medium - long trajectories (300 - 1000 points), long trajectories (1000 - 2000 points), and ultra - long trajectory sets (> 2000 points).

[0079] Step S2: Based on the distribution and change of the trajectory direction, roughly divide the trajectory. Calculate the direction of the trajectory to obtain a sequence of direction vectors of the trajectory. Based on the Kmeans clustering algorithm, identify potential trajectory inflection points, and divide the trajectory into a series of sub - trajectories at the potential inflection points to obtain a sequence of sub - trajectories.

[0080] 1. Calculate the connection direction between two consecutive points on a trajectory in sequence to obtain a sequence of direction vectors of the trajectory.

[0081] 2. Cluster the direction vectors of a trajectory based on the Kmeans clustering algorithm, where the number of clusters k will be determined according to the distribution of the direction vectors in eight direction intervals (eight core directions, 45° apart from each other), and the value range is [2, 8]; the distance metric involved in the algorithm uses cosine distance. The result of clustering is a sequence of cluster labels of the direction vector sequence, and the trajectory is divided at the label change to obtain a sequence of sub - trajectories.

[0082] Step S3: Calculate the merging threshold of the sub - trajectories based on cosine similarity, and merge some of the roughly divided sub - trajectories to obtain a sequence of merged sub - trajectories.

[0083] 1. Search for sub-trajectories with only two trajectory points in the sub-trajectory sequence as the sub-trajectories to be merged, and calculate the sub-trajectory merging threshold for them. The threshold will be adaptively calculated according to the position of the direction vector in the cluster (the maximum cosine distance between this direction vector and other vectors in the cluster to which it belongs is the merging threshold). If the vector is close to the cluster center, a smaller threshold will be obtained; otherwise, a larger threshold will be obtained.

[0084] 2. Calculate the dissimilarity in direction between the sub-trajectory and its adjacent sub-trajectories (the average cosine distance between this direction vector and all direction vectors in the previous or next sub-trajectory), and take the smaller value to compare with the threshold. If it is less than the threshold, merge the sub-trajectory with the adjacent sub-trajectory that is more similar in direction; otherwise, do not merge.

[0085] 3. After judging all the sub-trajectories to be merged, obtain the merged sub-trajectory sequence.

[0086] Step S4. Based on the improved silhouette coefficient, judge the conditions for the movement of the sub-trajectory boundary, move the boundaries of some sub-trajectories to obtain the updated sub-trajectory sequence. Adjacent sub-trajectories share a trajectory point, and extract such trajectory points as inflection points.

[0087] 1. Traverse each sub-trajectory in the sub-trajectory sequence. If a sub-trajectory contains only two trajectory points, skip it and do not consider it.

[0088] 2. If the current sub-trajectory is not the first one, execute step 3; otherwise, execute step 4.

[0089] 3. Obtain the direction vector corresponding to the first trajectory segment in the current sub-trajectory and calculate its silhouette coefficient. (The silhouette coefficient itself is an internal index for evaluating the clustering effect, and it is improved here for evaluating the trajectory partitioning effect; the main difference is that when calculating the external distance, considering the continuity characteristics of time series data, only the distance calculation with adjacent sub-trajectories is required.) If the silhouette coefficient is greater than 0, the boundary does not move, and execute step 4; otherwise, insert the trajectory segment corresponding to the current direction vector at the end of the previous sub-trajectory and remove it from the beginning of the current sub-trajectory, that is, move the boundary backward, and repeat step 3.

[0090] 4. If the current sub-trajectory is not the last one, execute step 5; otherwise, execute step 1 to traverse the next sub-trajectory.

[0091] 5. Obtain the direction vector corresponding to the last trajectory segment in the current sub-trajectory and calculate its silhouette coefficient. If the silhouette coefficient is greater than 0, the boundary does not move, and execute step 1 to traverse the next sub-trajectory; otherwise, insert the trajectory segment corresponding to the current direction vector at the beginning of the next sub-trajectory and remove it from the end of the current sub-trajectory, that is, move the boundary forward, and repeat step 5.

[0092] 6. After judging all sub-trajectories, a sub-trajectory sequence after boundary movement is obtained. Adjacent sub-trajectories share a trajectory point, and such trajectory points are extracted as inflection points.

[0093] Step S5. Retain the starting point, inflection points and ending point of the trajectory and delete other redundant trajectory points to obtain a compressed trajectory.

[0094] Following the above specific implementation steps, comparison results with two other direction-based trajectory compression methods are obtained on trajectory datasets at four different scale levels.

[0095] As Figure 2 shown, the running times of the three algorithms at the same compression ratio level are compared. Generally speaking, the running times of the three algorithms increase with the increase of the trajectory scale. Among them, the change range of the running time of the corner-based method is relatively small. From the graph, there is basically no obvious change, and the highest running time does not exceed 100 ms. Relatively speaking, the change range of the information entropy-based method in running time is relatively large, and it shows an exponential change trend with the increase of the trajectory scale, and the gap with the corner-based method becomes gradually obvious. The method of this embodiment is between the two, and the running time is lower than that of the information entropy-based method.

[0096] As Figure 3 shown, the compression ratios of the three algorithms at the same direction error level are compared. Generally speaking, the compression ratios of the three algorithms increase with the increase of the trajectory scale, and the highest reaches about 80%. Among them, the compression ratio using the method of this embodiment for compression is higher than that of the other two algorithms at any trajectory scale, indicating that the inflection point accuracy of the trajectory extracted by the method of this embodiment is relatively high, and the equivalent trajectory direction information can be retained with fewer trajectory points.

[0097] As Figure 4 shown, the direction errors of the three algorithms at the same compression ratio level are compared. Generally speaking, the direction errors generated by the three algorithms increase with the increase of the trajectory scale. The change of the direction error generated by the method of this embodiment fluctuates less, and can basically be kept below 5°. It has the lowest direction error at different trajectory scales. The direction errors generated by the other two algorithms fluctuate greatly. With the increase of the trajectory scale, the direction error gradually increases, and the gap with the method of this embodiment becomes gradually obvious. It once again shows that the inflection point accuracy of the trajectory extracted by the method of this embodiment is relatively high, and more trajectory direction information can be retained with equivalent trajectory points.

[0098] In this embodiment, to further prove the effectiveness of the proposed algorithm for identifying trajectory direction feature points, the inflection point recognition results of a representative trajectory in the short trajectory dataset, medium-long trajectory dataset, long trajectory dataset, and ultra-long trajectory dataset are respectively visualized, and the direction feature extraction method based on trajectory turning angle is used as a comparative algorithm for experimental analysis.

[0099] As Figure 5 shown, each row is the visualization result of a representative trajectory, from left to right are the method of the present invention, the method based on turning angle (taking a smaller threshold), and the method based on turning angle (taking a larger threshold). At the same time, the relevant metric indicators of the algorithm are evaluated as shown in Table 2.

[0100] Table 2

[0101]

[0102]

[0103] Based on Figure 5 the results shown in and Table 2, it can be clearly seen that the proposed algorithm can identify relatively accurate trajectory inflection points, without obvious omission of inflection points or excessive identification of redundant inflection points, and can produce less direction error while maintaining a high compression ratio. The inflection point recognition results of the method based on turning angle are greatly affected by the threshold. For the same trajectory, taking a larger or smaller threshold may result in the omission of some key inflection points or the identification of many unrepresentative inflection points, thus leading to inaccurate inflection point extraction results. And for different trajectories, different thresholds need to be selected to achieve better inflection point recognition effects respectively. In practical applications, due to the huge amount of trajectory datasets, it is almost impossible to set different parameter thresholds for different trajectories manually. Usually, through multiple parameter adjustments, a relatively good threshold is determined for the complete trajectory dataset, which will lead to the two extreme situations that the compression effect of some trajectories is better while that of some other trajectories is worse.

[0104] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0106] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0108] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

[0109] This patent is not limited to the above best implementation manner. Anyone can derive various other forms of the adaptive trajectory inflection point extraction and compression method and system based on partitioning under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage scope of this patent.

Claims

1. A partitioning-based adaptive trajectory inflection point extraction and compression method, characterized in that, it includes the following steps: Step S1: Obtain trajectory data, and after preprocessing, construct a basic trajectory data set; Step S2: Calculate the driving direction vectors of two adjacent trajectory points in sequence to obtain a sequence of direction vectors of the trajectory; cluster the sequence of direction vectors based on the Kmeans algorithm, roughly divide the trajectory into sub-trajectories, and obtain a sequence of sub-trajectories; Step S3: Search for sub-trajectories in the sequence of sub-trajectories that contain only two trajectory points as sub-trajectories to be merged, calculate the merging threshold of each sub-trajectory to be merged based on cosine similarity, and merge the sub-trajectories that need to be merged to obtain a merged sequence of sub-trajectories; Step S4: Based on the improved silhouette coefficient, judge the conditions for the boundary movement of the sub-trajectory, move the boundary of the sub-trajectory that needs to be moved, and obtain an updated sequence of sub-trajectories. Adjacent sub-trajectories share a trajectory point, which is used as the inflection point of the trajectory; Step S5: Extract the starting point, inflection point and ending point of the trajectory sequence, and the generated trajectory sequence is the compressed trajectory; Step S3 specifically includes the following steps: Step S31: Calculate a merging threshold for each sub-trajectory to be merged; the process description is as follows: calculate the cosine distance between the direction vector corresponding to the sub-trajectory and other direction vectors belonging to the same class, and take the maximum value as the merging threshold; the calculation formula is as follows: where x is described as the unique direction vector of the sub-trajectories to be merged, and x i is the other direction vector belonging to the same cluster as x, and the cosine distance is used as the similarity measure between vectors; Step S32: Calculate the average value of the cosine distances between the direction of each sub-trajectory to be merged and all direction vectors of adjacent sub-trajectories. If it is greater than the merging threshold, merge the sub-trajectory with the adjacent sub-trajectory with more similar direction to obtain a merged sequence of sub-trajectories; Step S4 specifically includes the following steps: Considering the time series characteristics of the trajectory data, calculate the silhouette coefficient at the boundary trajectory segment of the sub-trajectory, and determine whether to move the boundary of the sub-trajectory and the direction of the boundary movement according to the size of the silhouette coefficient; the calculation formula is as follows: Where: Wherein, the silhouette coefficient has a value between [-1, 1], the closer it is to 1, the relatively better the cohesion and separation degree of the direction vector are; conversely, the closer it is to -1, the relatively worse the cohesion and separation degree of the direction vector are; is to calculate the cohesion degree of the direction of the trajectory segment, that is, described as the current direction vector and the average value of the dissimilarity degree with other direction vectors within the same sub-trajectory; is to calculate the separation degree of the current direction vector, described as the current direction vector and the average value of the dissimilarity degree with all direction vectors within its directly adjacent sub-trajectory; During the boundary movement process, if the situation occurs that the current sub-trajectory only contains two trajectory points after the movement, judge the merging conditions of the sub-trajectory, and merge the sub-trajectories that need to be merged; Obtain the sequence of sub-trajectories after the boundary movement. Adjacent sub-trajectories share a trajectory point, and extract such trajectory points as the inflection points of the trajectory.

2. The partitioning-based adaptive trajectory inflection point extraction and compression method according to claim 1, characterized in that, the preprocessing in Step S1 includes: removing duplicate records and abnormal data and filling in missing data.

3. The partitioning-based adaptive trajectory inflection point extraction and compression method according to claim 1, characterized in that: Step S2 specifically includes the following steps: Step S21: Calculate the direction of each pair of consecutive trajectory points in sequence, and map a trajectory to a sequence of direction vectors; the calculation formula is as follows: where p describes a trajectory point in a trajectory, and lon and lat are the longitude and latitude of point p respectively; a sequence of direction vectors is described as Step S22: Divide the driving direction of the trajectory into eight core directions, with each direction spaced 45° apart. The eight core directions divide the driving direction into eight intervals; record the interval to which each direction vector in the direction vector sequence belongs, and count the total number of intervals to which they belong as the input parameter k required for the Kmeans clustering algorithm. The value range of k is [2, 8]. When all direction vectors are in the same interval, k is taken as 2; Step S23: Cluster the trajectory direction vectors, using the cosine distance for distance measurement; according to the clustering results, assign the same label to the direction vectors belonging to the same class, and perform trajectory division at the label change points to achieve a rough division of the trajectory, obtaining a sub-trajectory sequence.

4. A partition-based adaptive trajectory inflection point extraction and compression system, based on the method according to claim 1, characterized in that it includes: A rough trajectory division module, which is used to obtain a sequence of trajectory direction vectors by calculating the driving direction of the trajectory, and use the Kmeans clustering algorithm to cluster the vectors to discover the direction distribution and changes in the trajectory from a global perspective, so as to obtain a preliminary trajectory division result; A sub-trajectory merging module, which is used to adjust the result of the rough trajectory division from a local perspective, and appropriately merge some sub-trajectories based on cosine similarity to effectively reduce potential incorrect trajectory division results and improve the accuracy of trajectory division; A fine trajectory division module, which is used to adjust the result of the trajectory division from a local perspective again, and improve the silhouette coefficient to evaluate the result of the trajectory division from both the homogeneity within the sub-trajectory and the heterogeneity between the sub-trajectories, so as to further improve the accuracy of the trajectory division.