Trajectory data set management method
By preprocessing and compressing the trajectory dataset, calculating metadata, and establishing spatiotemporal range indexes in the memory key-value database, the storage and query efficiency problems of multiple trajectory datasets are solved, and efficient trajectory dataset management and query are achieved.
Patent Information
- Application Number
- CN202510280023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively manage and query multiple trajectory data sets, especially in terms of storage and query efficiency, and it is difficult to meet content-based query requirements.
By preprocessing and compressing the trajectory dataset, calculating metadata, and establishing spatiotemporal range indexes in the memory key-value database, efficient storage and querying of multiple trajectory datasets are achieved.
It realizes storing multiple trajectory data sets at a smaller space and time cost, and speeds up the efficiency of time range, spatial range and similar trajectory queries, helping users more easily discover and use the required trajectory data sets.
Smart Images

Figure CN120196697A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dataset management, and in particular to a management technology for trajectory datasets, including compressed storage, metadata management, and trajectory query methods. Background Art
[0002] Data is the basis of modern scientific research. The sharing of datasets (collections of homogeneous data organized in a certain way) can help numerous users study the theories of a certain field more objectively and deeply, and analyze the characteristics of actual systems efficiently and comprehensively. In the field of machine learning, datasets play a crucial role in training models, evaluating the generalization performance of models, etc. For data-driven tasks, the impact of the quality and quantity of datasets is more significant. However, due to the lack of infrastructure and incentives for dataset sharing, as well as effective management and complete documentation, existing shared datasets are difficult to use. The dataset search engine launched by Google aims to help users obtain various types of datasets more easily. However, such a general dataset search engine is difficult to meet the usage requirements of specific fields. The advent of the era of embodied intelligence has put forward higher requirements for the research and management of a large number of trajectory datasets. In scenarios such as autonomous driving and various robot applications, the effective management of trajectory datasets is a necessary prerequisite for data analysis, model prediction, and ensuring the safety of moving objects.
[0003] First of all, compared with general dataset management, the management of trajectory datasets faces challenges in many aspects. Since trajectory datasets contain unique spatio-temporal attributes, users of trajectory datasets often hope to obtain evaluation index results such as the sampling rate of the dataset that highlight its spatio-temporal characteristics, and general dataset search engines are inconvenient to provide such information. Therefore, specialized trajectory dataset metadata management technology is required. Secondly, the management of trajectory datasets not only needs to manage multiple trajectories, but also needs to solve the storage management problems of multiple trajectory datasets, and be able to provide query functions on multiple trajectory datasets, which is a task difficult to complete by trajectory management systems. In addition, existing dataset management technologies have not well solved the problem of content-based queries, and for trajectory dataset management, content-based query conditions are more common, and the increasing amount of massive trajectories further increases the difficulty of meeting this demand.
[0004] Chinese Patent Application No. 202311614352 discloses "A GNSS Trajectory Data Storage, Query Method and Database System". It creates a trajectory information summary table in the database, defines a TableInfo structure and creates a Map dictionary; extracts keywords from user trajectory data and generates a random number; stores the data in different TableInfo structures according to the random number; and converts the user trajectory data in the TableInfo structure into a batch insert statement and stores it in the corresponding sub-table for storing trajectory information, which can make the distribution of massive GNSS trajectory data uniform and improve the data storage speed. When querying data, it groups the data according to the number of threads and improves the data query efficiency through the multi-threaded multi-table joint query method. It does not minimize the storage space from the perspective of compressing trajectories, and it targets a single trajectory dataset rather than the storage and query of multiple trajectory datasets.
[0005] Chinese Patent Application No. 202410933671 discloses "A Trajectory Data Processing Method and Device". It receives natural language query information for a non-relational trajectory storage database and converts the natural language query information into a target query request formed by a database query statement; extracts query keywords containing a time range and / or a space range from the target query request, and extracts corresponding target trajectories for a single object or multiple objects from the non-relational trajectory storage database based on a preset multi-level index. This invention does not provide a query method for multiple trajectory datasets, and it does not design a query method for similar trajectories.
[0006] Chinese Patent Application No. 202410958872 discloses "A Trajectory Data Storage Method and System for Edge Environments". It constructs a lightweight, network-aware general framework for the analysis and storage of edge trajectories, adopts a storage-computation separation architecture, divides the nodes on the edge into computing nodes and storage nodes, where the computing nodes are responsible for lightweight tasks such as time alignment of trajectories, road network matching, inverted index construction, and customized compression. A co-flow control algorithm based on credit values is used for data transmission and communication between different computing nodes. The storage nodes form a virtual connection ring using a decentralized distributed consistent hashing algorithm and store data using a log-structured merge tree structure. It mainly focuses on the trajectory storage in edge environments and does not implement a query function for trajectories. Summary of the Invention
[0007] Objective of the Invention: The problem to be solved by the present invention is how to store multiple trajectory data sets at a relatively small space cost, and on this basis, help users screen out data sets that better meet their metadata requirements and content requirements at a relatively small time cost. Further, the objective of the present invention also lies in designing a time and space index and corresponding search methods to accelerate the time range, space range query, and Top-k similar trajectory query of the trajectory data sets.
[0008] Technical Solution: To achieve the above objective of the invention, the technical solution adopted by the present invention is a method for managing trajectory data sets, including the following steps:
[0009] (1.1) Preprocess and compress the trajectory data sets to obtain compressed trajectory data sets;
[0010] (1.2) Calculate the metadata of the compressed trajectory data sets according to the designed evaluation metrics;
[0011] (1.3) Establish a spatio-temporal range index and store multiple compressed trajectory data sets in an in-memory key-value database;
[0012] (1.4) Search and display the compressed trajectory data sets based on a search engine according to the query conditions.
[0013] Further, the step (1.1) includes the following steps:
[0014] (2.1) First, specify the normal speed threshold for the corresponding trajectory according to the type of moving object, remove the noise trajectory points exceeding the normal speed threshold, and then for the missing timestamps or trajectory points, use the method of linear interpolation to fill them;
[0015] (2.2) Based on the trajectory data sets after preprocessing obtained in the step (2.1), segment each trajectory in the trajectory data sets according to the positional relationship between the trajectory points and the compression distance threshold, retain the start point and end point of each segment of the trajectory, and remove some intermediate points according to the compression distance threshold to form compressed trajectory data sets.
[0016] Further, the step (1.2) includes the following steps:
[0017] (3.1) Based on the step (1.1), calculate the time range \(R\) T and the space range \(R\) S of each compressed trajectory data set, as well as the time interval sequence and space interval sequence of each trajectory;
[0018] (3.2) Based on the step (3.1), for each compressed trajectory data set, calculate the average time interval and average space interval to obtain two types of metadata, namely the average time sampling rate and the average space sampling rate;
[0019] (3.3) Based on the step (3.1), for each compressed trajectory dataset, calculate the variance of the time interval sequence and the variance of the space interval sequence for each trajectory, and obtain two types of metadata: time sampling sparsity and space sampling sparsity.
[0020] Further, the step (1.3) includes the following steps:
[0021] (4.1) Use quadtree node encoding to preliminarily represent the relative position of each trajectory within the spatial range R S of the compressed trajectory dataset to which it belongs; further refine the description of the spatial relative position of each trajectory using position encoding, and mark all subspaces of the parent space covered by the trajectory; combine the quadtree node encoding of the trajectory and the trajectory ID as the key, and store the trajectory point sequence, position encoding, the start and end times corresponding to the trajectory, and the minimum bounding rectangle as the value into the in-memory key-value database to establish a spatial range index;
[0022] (4.2) According to the time range R T of the compressed trajectory dataset and the set fixed time period length, calculate the serial number of the time period, which is called the segment number; within the time period, encode different sub-segments using binary tree encoding and appending 0 or 1 at the end, which is called the sub-segment code; combine the segment number, sub-segment code, and trajectory ID as the key, and the time range R Ti of the i-th trajectory as the value, and store it into the in-memory key-value database to establish a time range index.
[0023] Further, the step (1.4) includes the following steps:
[0024] (5.1) Based on the metadata of the compressed trajectory dataset obtained in the steps (3.2) and (3.3), filter out the compressed trajectory datasets that meet the metadata query conditions;
[0025] (5.2) Based on the compressed trajectory datasets obtained in the step (5.1), further screen out one by one the compressed trajectory datasets related to the content query conditions;
[0026] (5.3) Count the number of trajectories that meet the query request or the sum of the similarities of the top-k similar trajectories that meet the query request in each compressed trajectory dataset obtained in the step (5.2), and then sort the compressed trajectory datasets in descending order as the query result.
[0027] Further, the step (5.2) includes the following steps:
[0028] (6.1) For the content query condition of querying the time range R Tq , first calculate the time range R T of the compressed trajectory dataset that overlaps with the query time range RTq All time periods with intersections; the trajectories in the middle time period all meet the query request; for the trajectories in the start time period or the end time period, further according to the time range R of the i-th trajectory among them Ti and the query time range R Tq Verify one by one, and the trajectories with overlaps meet the query request;
[0029] (6.2) For the content query condition of querying the spatial range R Sq When, assume that the query spatial range R Sq is a trajectory T containing four trajectory points, and the four trajectory points respectively correspond to the four vertices of the query spatial range R Sq According to the quadtree node encoding and position encoding of the trajectory T within the spatial range R S in the compressed trajectory data set, filter out the trajectories that may intersect with the trajectory T as candidate trajectories, and then eliminate the trajectories that cannot intersect with the query spatial range R according to the minimum bounding rectangle of the candidate trajectories Sq Finally, verify the remaining candidate trajectories one by one according to the longitude and latitude of the trajectory points to obtain all the trajectories that meet the query spatial range R Sq ;
[0030] (6.3) For the Top-k similar trajectory query, based on the similarity distance threshold, use the best-first heuristic method to quickly locate the spatial range that may contain similar trajectories, and then obtain the k trajectories that are most similar to the query trajectory.
[0031] Beneficial effects: (1) The present invention can effectively help users screen a trajectory data set that better meets their needs based on the metadata and content of the trajectory data set, and more conveniently discover and use the required trajectory data set; (2) The present invention can reduce the storage cost of the trajectory data set by compressing the trajectories and designing a spatio-temporal index, and can achieve a balance between the storage space and the query accuracy; (3) By pre-storing multiple compressed trajectory data sets and the designed spatio-temporal index and their corresponding query methods, the present invention can accelerate the search efficiency of the trajectory data set in terms of time range query, spatial range query, and similar trajectory query. By adopting the technical solution of the present invention, engineers can relatively easily implement related software. Brief Description of the Drawings
[0032] Figure 1 is the overall processing flow chart of the present invention;
[0033] Figure 2 is the schematic diagram of quadtree node encoding;
[0034] Figure 3 is the schematic diagram of position encoding;
[0035] Figure 4 is the schematic diagram of spatial range query. Detailed implementation manners
[0036] The present invention will be further clarified below in conjunction with the accompanying drawings and specific implementation examples. It should be understood that these implementation examples are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications made by those skilled in the art to the present invention fall within the scope defined by the appended claims of this application.
[0037] The present invention provides a method for managing a trajectory dataset. Based on data preprocessing and compression technologies, multiple compressed trajectory datasets are stored in an in-memory key-value database, and efficient management of the trajectory dataset is achieved by using metadata management and encoding technologies, facilitating the discovery and use of the trajectory dataset and giving play to the data value.
[0038] The complete process of the present invention is as Figure 1 shown, including four major parts: collecting the trajectory dataset and performing preprocessing and compression; designing metrics, calculating metadata, and storing them in a search engine; establishing a spatio-temporal range index using encoding technology and storing the compressed trajectory dataset; returning the trajectory dataset that meets the requirements according to the query conditions input by the user.
[0039] The implementation manners of the present invention will be illustrated by examples in modules below:
[0040] 1. Preprocess the trajectory dataset to be managed and then compress it.
[0041] 1.1 Preprocess the trajectory dataset to be pre-stored. First, check whether each trajectory in the trajectory dataset has missing data, such as missing spatial attribute values at certain time points. For the missing data, linear interpolation is used for supplementation, and the position or attribute of the missing point is estimated based on the information of the front and back data points. Then, redundant fields are removed, and only the trajectory ID and the sequence of trajectory points (with timestamps) are retained. Then, abnormal data is identified and removed by moving objects according to their normal speed thresholds (for example, the speed of a bicycle should be below 40 km / h), and finally, reasonable trajectories are retained.
[0042] 1.2 Compress each trajectory in the trajectory dataset, and then complete the compression of the trajectory dataset. Assume that a trajectory has n sampling points, p i represents the i-th sampling point among them, represents the vector from the 1st sampling point to the i-th sampling point, represents the vector from the 1st sampling point to the n-th sampling point, represents the vector from the i-th sampling point to the n-th sampling point, d i represents the compression distance of the i-th sampling point. Then when and , d i is p iThe perpendicular distance to the line p1p n ; otherwise, d i = min(distance(p1, p i ), distance(p i , p n ))), where distance represents the Euclidean distance between two points. The following is an example to illustrate the compression process. For example, assuming the pre-set compression distance threshold is 1, the original sampling point sequence of a certain trajectory is {p1, p2, p3, p4, p5}, where p1: (0, 0), p2: (1, 1), p3: (3, 0), p4: (6, 6), p5: (5, 0), the starting point of the trajectory is p1, the ending point is p5, and {p2, p3, p4} are intermediate points. Among them, d2 and d3 are the perpendicular distances from p2 and p3 to the line where p1 and p5 are located, that is, 1 and 0, and d4 = min(distance(p4, p1), distance(p4, p5)) = 1.4. Since d4 is greater than the pre-set compression distance threshold 1, p4 is set as the segmentation point, and the whole trajectory is divided into two segments. Then, the above operations are performed on these two segments of the trajectory respectively.
[0043] At this time, the second segment of the trajectory has only two points p4 and p5, and no further operation is required. The starting point of the first segment of the trajectory is p1, the ending point is p4, and {p2, p3} are intermediate points. Among them, d2 and d3 are the perpendicular distances from p2 and p3 to the line where p1 and p4 are located, that is, 0 and 2.1. Since d3 is greater than the pre-set compression distance threshold 1, p3 is set as the segmentation point, and the first segment of the trajectory is divided into two segments again.
[0044] Now the second segment of the trajectory has only two points p3 and p4, and no further operation is required. The starting point of the first segment of the trajectory is p1, the ending point is p3, p2 is the intermediate point, and d2 is the perpendicular distance from p2 to the line where p1 and p3 are located, that is, 1. Since d2 does not exceed the compression distance threshold, p2 can be removed. The finally remaining trajectory points form the compressed trajectory {p1, p3, p4, p5}.
[0045] 2. Calculate the metadata of the compressed trajectory dataset.
[0046] Suppose there are the following several trajectory datasets to be managed:
[0047] Shanghai - mobike1: The first Shanghai bike trajectory dataset
[0048] Shanghai - mobike2: The second Shanghai bike trajectory dataset
[0049] Shanghai - mobike3: The third Shanghai bike trajectory dataset
[0050] Shanghai-taxi: The first taxi trajectory dataset in Shanghai
[0051] Beijing-mobike: The first bike trajectory dataset in Beijing
[0052] Beijing-taxi: The second taxi trajectory dataset in Beijing
[0053] 2.1 Traverse each trajectory in the compressed trajectory datasets, and calculate the metadata corresponding to each trajectory dataset that can reflect its time attribute (subscript T) and space attribute (subscript S). The time range R T and the space range R S are obtained by traversing all the data in the compressed trajectory datasets.
[0054] RATE T and RATE S represent the time sampling rate and space sampling rate of the trajectory respectively. For trajectory T, their calculation formulas are as follows:
[0055]
[0056] where t i refers to the time of the i-th sampling point in the trajectory, p i refers to the position of the i-th sampling point in the trajectory, distance(p i , p i-1 ) refers to the Euclidean distance between the i-th sampling point and the (i - 1)-th sampling point in the trajectory.
[0057] SPARSITY T (set) and SPARSITY S (set) are two operations used to calculate the sparsity of trajectory points to reflect whether the distribution of the trajectory dataset is uniform. The sparsity SPARSITY T (set) in terms of time is represented by the standard deviation of the time interval sequence of trajectory points, and the sparsity SPARSITY s (set) in terms of space is represented by the standard deviation of the spatial distance sequence of trajectory points. The calculation formulas are as follows:
[0058]
[0059] where S i represents the Euclidean distance between the i-th sampling point and the (i - 1)-th sampling point, and T i represents the time difference between the i-th sampling point and the (i - 1)-th sampling point.
[0060] 2.2 As shown in Table 1, store these calculated metadata together with information such as the location city and moving objects in the search engine to facilitate subsequent query layer to search for the trajectory dataset based on the metadata.
[0061] Table 1 Example of Metadata of Trajectory Dataset
[0062]
[0063] 3. Design time range query index and space range query index to store the trajectory dataset.
[0064] 3.1 Assume that the compressed trajectory dataset has 3 trajectories as shown in Figure 2 . Calculate the quadtree node encoding of each trajectory according to the spatial range of the entire trajectory dataset, and obtain the location encoding by referring to Figure 3 . The final result is shown in Table 2.
[0065] Table 2 Example of Trajectory Encoding
[0066] Trajectory ID Quadtree node encoding Location encoding T1 1 4 T2 12 9 T3 100 6
[0067] 3.2 Store the trajectory dataset shown in Table 2 in the in-memory key-value database in the form of Key-Value, as shown in Table 3, where PL i refers to the sequence of trajectory points of the i-th trajectory, t si and t ei refer to the start time and end time of the i-th trajectory respectively, and mbr i refers to the minimum bounding rectangle of the i-th trajectory.
[0068] Table 3 Storage of Trajectory Dataset
[0069]
[0070] 4. Support users to search for the trajectory dataset based on metadata and search for the trajectory dataset based on content.
[0071] 4.1 Based on the trajectory dataset in the example given in 1, the user inputs the query conditions: location city: Shanghai, moving object: bicycle, time sampling rate: less than 0.5 times / s. Convert such requirements into the parameters of the corresponding interface function provided by the search engine, and then return two trajectory datasets, Shanghai-mobike1 and Shanghai-mobike2, that meet the query request.
[0072] 4.2 After the user obtains two datasets, Shanghai-mobike1 and Shanghai-mobike2, that meet their metadata requirements, further screen the datasets according to the content requirements. The content query here involves time range query, space range query, and Top-k similar trajectory query.
[0073] 4.2.1 For time range query, first calculate the time periods B1, B2... B Tq : [t s , t e that intersect with the query time range R n-1 , B n , Since the query time range completely covers the middle time period, all the trajectories in the dataset that intersect with the middle time period meet the query request. For the start time period and the end time period, calculate the in-segment index of the query time range in this time period, find the candidate trajectories according to the in-segment index, and finally verify whether the time range of the candidate trajectories overlaps with the query time range. The trajectories with overlapping time ranges meet the query request.
[0074] 4.2.2 For spatial range query, based on the Shanghai-mobike1 trajectory dataset, the query spatial range R Sq as Figure 4 shown. First, assume R Sq as a trajectory containing four trajectory points, and the four trajectory points correspond to the four vertices of R Sq respectively. Based on this, calculate the spatial index encoding index and position encoding PC of R Sq within the entire trajectory dataset spatial range, and then use the (index, PC) combination to represent the approximate area of R Sq in the entire trajectory dataset space to the greatest extent. Figure 4 The query range R Sq in
[0075] Table 4 Spatial index encoding and position encoding examples corresponding to trajectories
[0076] Trajectory (index, PC) T1 (1,6) T2 (10,4) T3 (101,4) T4 (10,1)
[0077] Thus, the three trajectories T1, T2, and T3 can be filtered out. Because although the node where trajectory T1 is located is the parent node of R Sq , but the PC of T1 is 6, indicating that it has no trajectory points in the lower left subspace, so there cannot be trajectory points within R Sq in trajectory T1. Although the index of trajectory T2 is the same as that of R Sq , but its PC is 4, indicating that it does not contain the trajectory points in the left half subspace of the lower left subspace, so there are no trajectory points within R Sq in it. The index of trajectory T3 is 101, indicating that all its trajectory points are located in the lower right subspace of the subspace corresponding to index "10", while R SqThe position encoding is 1, indicating that it is related to R Sq The right half subspace of the lower left subspace does not intersect with R, so T3 does not contain R Sq The trajectory points within it. And T4 contains trajectory points within the query range R Sq Among them, the query result is the trajectory T4. In the same way as above, the trajectories that meet the query condition R in Shanghai-mobike2 can also be obtained Sq of the trajectories.
[0078] 4.2.3 For the Top-k similar trajectory query. Maintain three priority queues tq, ipq, iq and the similarity distance threshold maxDist. The candidate trajectories are stored in tq. The candidate (quad-tree node encoding, position encoding) combinations are stored in ipq, and each combination represents a refined spatial range. The candidate quad-tree node encodings are stored in iq, and each encoding represents a spatial range. maxDist is initialized to a very large value. When k candidate trajectories have been stored in tq, maxDist is updated to the similarity of the least similar trajectory among the current k candidate trajectories to the query trajectory. When iq is initialized, it stores the encoding of the root node of the quad-tree representing the spatial range of the compressed trajectory dataset.
[0079] The first-layer loop processing: When the queue iq is not empty, take out an element i from it. Judge whether the queue ipq is empty, or whether the distance between the first element of the ipq queue and the query trajectory is greater than the distance between the element i and the query trajectory. If the above conditions are met, further judge whether the distance between the element i and the query trajectory is greater than maxDist. If it is greater, terminate the loop; if it is not greater, add the refined space of the element i to the queue ipq, and add all the sub-spaces of the element i to the queue iq.
[0080] The second-layer loop processing: When the queue ipq is not empty and the distance between the elements in the ipq and the query trajectory is less than the distance between the element i and the query trajectory, take out an element ip from the queue ipq. Judge whether the distance between the element ip and the query trajectory is greater than maxDist. If it is greater, output the current queue tq as the result. If it is not greater, traverse each trajectory t in the compressed trajectory dataset within the spatial range represented by the element ip. For each trajectory t, calculate its distance from the query trajectory. If the distance is less than maxDist, add the trajectory t to the queue tq and update maxDist.
[0081] Finally, the k trajectories in the compressed trajectory dataset that are most similar to the query trajectory are contained in the queue tq, and they are output as the result.
[0082] 4.3 For time range queries and space range queries, the relevance is defined as the number of trajectories in the trajectory dataset that meet the query conditions. For Top-k similar trajectory queries, the relevance is defined as the sum of the similarities of the k returned trajectories. After calculating the relevance, the dataset is sorted in descending order of relevance and returned to the user so that the user can filter the trajectory dataset that better meets their content requirements.
Claims
1. A trajectory data set management method, characterized in that: The steps include: (1.1) Preprocessing and compressing the trajectory data set to obtain a compressed trajectory data set; (1.2) Calculate the metadata of the compressed trajectory dataset according to the designed evaluation index; (1.3) Establish a spatiotemporal range index and store multiple compressed trajectory data sets in an in-memory key-value database; (1.4) Based on the query conditions, the compressed trajectory dataset is searched and displayed based on the search engine.
2. The trajectory data set management method according to claim 1, characterized in that: The step (1.1) comprises the following steps: (2.1) First, the normal speed threshold of the corresponding trajectory is specified according to the type of moving object, and the noise trajectory points that exceed the normal speed threshold are removed. Then, the missing timestamps or trajectory points are filled by linear interpolation; (2.2) Based on the preprocessed trajectory data set obtained in step (2.1), each trajectory in the trajectory data set is segmented according to the positional relationship between trajectory points and the compression distance threshold, the starting point and the end point of each trajectory segment are retained, and some intermediate points are removed according to the compression distance threshold to form a compressed trajectory data set.
3. The trajectory data set management method according to claim 1, characterized in that: The step (1.2) comprises the following steps: (3.1) Based on step (1.1), calculate the time range R of each compressed trajectory data set T and spatial range R S , as well as the time interval sequence and spatial interval sequence of each trajectory; (3.2) Based on step (3.1), for each compressed trajectory data set, the average time interval and the average spatial interval are calculated to obtain two kinds of metadata: the average time sampling rate and the average spatial sampling rate; (3.3) Based on the step (3.1), for each compressed trajectory data set, the time interval series variance and the space interval series variance of each trajectory are calculated to obtain two kinds of metadata: time sampling sparsity and space sampling sparsity.
4. The trajectory data set management method according to claim 3, characterized in that: The step (1.3) comprises the following steps: (4.1) Quadtree node encoding is used to preliminarily represent the spatial range R of each trajectory in the compressed trajectory dataset to which it belongs. S The relative position within the space; further use the position code to refine the spatial relative position of each trajectory, and mark all subspaces of the parent space covered by the trajectory; the quadtree node code of the trajectory and the trajectory ID are combined as the key, and the trajectory point sequence, position code, start and end time and minimum enclosing rectangle corresponding to the trajectory are used as the value, and stored in the memory key-value database to establish a spatial range index; (4.2) According to the time range R of the compressed trajectory dataset T The serial number of the time period is calculated based on the fixed time period length, which is called the segment number. Within the time period, different sub-segments are encoded by using binary tree coding and adding 0 or 1 at the end, which is called the sub-segment code. The segment number, sub-segment code and trajectory ID are combined as the key, and the time range R of the i-th trajectory is Ti As the value, it is stored in the in-memory key-value database to create a time range index.
5. The trajectory data set management method according to claim 3, characterized in that: The step (1.4) comprises the following steps: (5.1) Based on the metadata of the compressed trajectory dataset obtained in steps (3.2) and (3.3), filter out the compressed trajectory dataset that meets the metadata query condition; (5.2) Based on the compressed trajectory data set obtained in step (5.1), further filtering out the compressed trajectory data sets related to the content query condition one by one; (5.3) Counting the number of trajectories that meet the query request in each compressed trajectory data set obtained in step (5.2) or the sum of the similarities of the top-k similar trajectories that meet the query request, and then arranging each compressed trajectory data set in descending order as the query result.
6. The trajectory data set management method according to claim 5, characterized in that: The step (5.2) comprises the following steps: (6.1) When the content query condition is the query time range R Tq When the time range R of the compressed trajectory dataset is first calculated T The query time range R Tq All time periods with intersections; the trajectories in the middle time period all meet the query request; for the trajectories in the start time period or the end time period, further according to the time range R of the i-th trajectory Ti and query time range R Tq Verify one by one that there are overlapping tracks that meet the query request; (6.2) When the content query condition is the query space range R Sq When assuming that the query space range is R Sq is a trajectory T containing four trajectory points, and the four trajectory points correspond to the query space range R Sq The four vertices of trajectory T are in the spatial range R of the compressed trajectory dataset. S The quadtree node encoding and position encoding in the query space are used to select the trajectories that may intersect with trajectory T as candidate trajectories, and then the candidate trajectories are eliminated according to the minimum bounding rectangle of the query space range R. Sq There can be no intersection of trajectories; finally, the remaining candidate trajectories are verified one by one according to the longitude and latitude of the trajectory points, and all the trajectories that meet the query space range R are obtained. Sq trajectories; (6.3) For Top-k similar trajectory queries, based on the similarity distance threshold, the best-first heuristic method is used to quickly locate the spatial range that may contain similar trajectories, and then obtain the k trajectories that are most similar to the query trajectory.
Citation Information
Patent Citations
GNSS (Global Navigation Satellite System) trajectory data storage and query method and database system
CN117648391A
Track data processing method, device and equipment
CN118470616A
Track data storage method and system oriented to edge environment
CN118509434A
Spatial-temporal trajectory indexing and query processing method and device, equipment and medium
CN114117260A
Large-scale trajectory data space-time accompanying person query method and system
CN115658737A