Track data-oriented storage and index construction method

By using S2 spatial coding and binary time coding methods in trajectory data processing, and combining GeoMesa and HBase to build a storage model, the problem of low efficiency in large-scale trajectory data storage and query is solved, and efficient spatio-temporal query and system stability are achieved.

CN119938802AInactive Publication Date: 2025-05-06KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411892297.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art processes large-scale and frequently updated trajectory data, there are problems such as large storage overhead and low query efficiency, especially in the high concurrency and rapid retrieval requirements of massive trajectory data.

Method used

S2 spatial encoding and binary time encoding methods are adopted, and the underlying storage model is built in combination with GeoMesa and HBase, the index structure of partition ID, HID, TID, and FID is designed, and the S2 index is constructed and stored in the distributed database HBase, and the space-time query strategy is optimized.

Benefits of technology

It effectively improves the spatial proximity of trajectory data and the joint query capabilities of space-time, provides high-efficiency space-time query support. The system maintains the stability of the index structure when high-frequency data is updated, and is suitable for application scenarios such as traffic monitoring and geographic information systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938802A_ABST
    Figure CN119938802A_ABST
Patent Text Reader

Abstract

The invention discloses a track data-oriented storage and index construction method, which belongs to the technical field of big data management and indexing, and comprises the following steps of: 1) preprocessing original track data; 2) S2 space coding; 3) time coding; 4) designing an index table; (5) a GeoMesa-HBase trajectory data storage model is constructed; 6) constructing an index table; and 7) spatio-temporal query of an optimization strategy. The method has the beneficial effects that by introducing the S2 space coding and time coding method, the space proximity and the space-time joint query capability of the trajectory data are improved; by designing index structures of partition IDs, HIDs, TIDs and FIDs, trajectory data are uniformly distributed in distributed nodes, and high-efficiency space-time query support is provided; a bottom-layer storage model is constructed by utilizing GeoMesa and HBase, so that the stability of an index structure is kept when the system updates high-frequency data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data management and indexing technology, and in particular to a storage and indexing method for trajectory data. Background Art

[0002] The development of location services and navigation technologies has led to an explosive growth in the generation and accumulation of trajectory data, especially in the fields of traffic monitoring and geographic information systems (GIS). However, the diversity and complexity of trajectory data, including the heterogeneity of temporal and spatial distribution and the need for high-frequency updates, make traditional storage and query solutions face many challenges when processing large-scale trajectory data.

[0003] Existing relational databases have the limitations of high storage overhead and low query efficiency in storing massive trajectory data, making it difficult to meet the needs of trajectory data for high concurrency and fast retrieval. On the other hand, non-relational databases (NoSQL) have obvious advantages in data storage due to their high scalability, but they usually lack efficient spatiotemporal indexing mechanisms and cannot effectively solve the problem of efficient query of trajectory data in spatial and temporal dimensions.

[0004] To improve the indexing efficiency of trajectory data, the current mainstream methods include tree structure index structures (such as R-tree, quadtree), tree structure variants, and space filling curves. Tree structure indexes can support the management of spatiotemporal data to a certain extent, but there are large storage overheads and efficiency bottlenecks when massive data are dynamically updated and complex queries are performed. In contrast, space filling curve technology can significantly reduce index maintenance costs while maintaining data spatial proximity by mapping multidimensional data to one-dimensional data. Among space filling curves, although the Z curve is widely used for trajectory data indexing, its spatial clustering effect is relatively poor, and spatial clustering has a certain impact on the distribution and query of spatiotemporal data. Although the existing technology has optimized the storage and indexing efficiency of trajectory data to a certain extent, there is still room for improvement when dealing with large-scale and frequently updated trajectory data. Summary of the invention

[0005] To solve the above problems, the present invention provides a storage and index construction method for trajectory data, which is implemented by the following technical solutions.

[0006] A storage and index construction method for trajectory data, the method comprising the following steps:

[0007] 1) Preprocessing of original trajectory data:

[0008] Get the GPS trajectory data of the taxi, which includes spatial coordinate information, time information and vehicle ID;

[0009] Clean the trajectory data and retain the valid trajectory data;

[0010] Unified data format;

[0011] 2) S2 spatial coding:

[0012] Perform S2 spatial encoding on the spatial coordinate information in the trajectory data processed in step 1) to generate a one-dimensional index value to obtain a HID;

[0013] 3) Time coding:

[0014] The time information in the trajectory data processed in step 1) is converted into hexadecimal and combined to obtain TID;

[0015] 4) Index table design:

[0016] Design the Rowkey structure in the index table, combine the partition ID, HID, TID and vehicle identifier to generate the Rowkey, and ensure that each piece of data has a unique identifier;

[0017] The partition ID is the partition key, which divides and stores the trajectory data by day; FID is the vehicle ID;

[0018] 5) Build GeoMesa-HBase trajectory data storage model:

[0019] Combine the horizontal scalability of HBase with the GeoMesa spatiotemporal indexing tool to build an underlying data storage model;

[0020] 6) Build index table:

[0021] Based on the index table designed in step 4), use GeoMesa to build an S2 index and store the data in the distributed database HBase;

[0022] 7) Spatiotemporal query optimization strategy:

[0023] Set query parameters according to query conditions;

[0024] Construct a query object;

[0025] Determine the scope of the query;

[0026] Retrieve data and return query results.

[0027] Preferably, in the step 1):

[0028] The acquired trajectory data information is classified by day, and the acquisition frequency is 30s / time. The trajectory data information also includes whether there are passengers, speed and trajectory ID;

[0029] Data cleaning includes removing missing values, removing redundant values, removing data outside the study area, removing values ​​with abnormal passenger instantaneous states, and processing trajectory drift data;

[0030] Unifying the data format includes unifying the sorting of data and unifying the data format.

[0031] Preferably, in the step 2), the S2 space coding includes spherical coordinate transformation, cube projection, projection correction, gridding processing and Hilbert curve coding processing performed in sequence.

[0032] Preferably, the specific operations of various processing of the S2 space coding are as follows:

[0033] Spherical coordinate transformation converts the longitude and latitude data (lat, lng) in the spatial coordinate information into spatial rectangular coordinates (x, y, z). The calculation formula is:

[0034]

[0035] Assume that the radius of the earth is r = 1, where θ represents latitude, φ represents longitude, and the range of x, y, and z is in the interval [-1, 1];

[0036] Cube projection, construct a circumscribed cube with a side length of 1 outside the earth, realize the projection from the center of the sphere to the six faces of the cube, and obtain (f, u, v), where f is the face identifier, and (u, v) corresponds to the projection of the plane rectangular coordinates (x, y);

[0037] Projection correction, using quadratic transformation to adjust the projection coordinates (u, v), transforming (f, u, v) into (f, s, t), the calculation formula is:

[0038]

[0039] Where s, t are the corrected coordinates of u, v, and the range is [0, 1];

[0040] Grid processing: divide the two-dimensional plane coordinates (u, v) into quadtree grid units according to the set level N and convert them into grid coordinates (i, j), where i and j represent the row and column numbers of the grid respectively, and the value range of i and j is [0,2 N -1];

[0041] Hilbert curve encoding processing uses the Hilbert curve to encode the grid coordinates and combines them with the surface identifier to form a CellID, which is finally converted into a decimal system to obtain a HID.

[0042] Preferably, in step 6),

[0043] Before building the index, use the GeoMesa-HBase model API to traverse the main table data in the underlying storage;

[0044] Then build the index table. The value in the index table is the row key value in the main table.

[0045] The index table is built and stored in HBase, which is implemented through the GeoTools DataStore interface;

[0046] The mapping relationship between the index table and the main data table is constructed according to the designed Rowkey, and the data in the main table is located through the index table.

[0047] Preferably, in step 7),

[0048] When setting query parameters according to query conditions, determine the spatial range (Lat1, Lng1, Lat2, Lng2), time range (Stime, Etime), level N and table structure information, where Stime and Etime represent the query start time range and end time range, and level N corresponds to the spatial division in spatial coding;

[0049] When constructing a query object, use the query parameters as the query input and use the ECQL syntax to filter and decompose the query conditions;

[0050] When determining the query scope, the query scope of the HBase database is determined by the index of the query object;

[0051] Retrieve data and return query results.

[0052] The beneficial effects of the present invention are as follows: the present invention effectively improves the spatial proximity and spatiotemporal joint query capability of trajectory data by introducing S2 spatial coding and binary-like time coding methods; secondly, by designing the index structure of partition ID, HID, TID, and FID, the trajectory data is evenly distributed in distributed nodes, providing more efficient spatiotemporal query support; GeoMesa and HBase are used to build the underlying storage model, so that the system maintains the stability of the index structure when high-frequency data is updated; overall, the present invention has good effects in large-scale trajectory data storage, real-time analysis, system scalability and efficient query, and is suitable for application scenarios such as traffic monitoring and geographic information systems that have strict requirements on real-time data processing and efficient query, and has significant social and economic value. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the description of the specific implementation methods will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0054] Figure 1 : The process of preprocessing the original trajectory data;

[0055] Figure 2 : Schematic diagram of the comparison of the original trajectory data before and after preprocessing;

[0056] Figure 3 :S2 spatial coding process diagram;

[0057] Figure 4 : Schematic diagram of the principle of index table construction;

[0058] Figure 5 : Flowchart of spatiotemporal query. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] like Figure 1 As shown, a storage and index construction method for trajectory data includes the following steps:

[0061] 1) Preprocessing of original trajectory data:

[0062] Get the taxi GPS trajectory data information, which includes spatial coordinate information, time information and vehicle ID.

[0063] The acquired trajectory data information is classified by day, with an acquisition frequency of 30s / time. The trajectory data information also includes whether there are passengers, speed and trajectory ID. The daily data size is about 2-3GB.

[0064] Clean the trajectory data and retain the valid trajectory data.

[0065] In order to improve data quality and data integrity, data cleaning is performed on the data. Data cleaning includes:

[0066] Remove missing values, that is, remove records with null values;

[0067] Remove redundant values, that is, remove repeated or unnecessary information;

[0068] Remove data outside the study area and use shp or geojson files to remove data outside the study area;

[0069] Remove abnormal values ​​of passengers' instantaneous states;

[0070] As well as trajectory drift data processing, since vehicle trajectory data is generally affected by multiple factors during the collection process, resulting in deviations between the collected data and the actual driving path of the vehicle, causing the recorded trajectory information to be inconsistent with the actual situation, trajectory drift data processing is required.

[0071] Figure 2 This is a schematic diagram for comparing data before and after data preprocessing.

[0072] Unify the data format to provide a consistent data structure for subsequent indexing and storage.

[0073] The unified data format includes unified data sorting and unified data format, that is, each track data is arranged in a fixed order, and the information in each track data exists in a fixed format, such as the format of time information is unified as "YYYY-MM-DDHH:mm:ss".

[0074] 2) S2 spatial coding:

[0075] The spatial coordinate information in the trajectory data processed in step 1) is subjected to S2 spatial encoding to generate a one-dimensional index value to obtain HID. The S2 spatial encoding process is as follows: Figure 3 shown.

[0076] S2 space coding includes spherical coordinate transformation, cube projection, projection correction, gridding and Hilbert curve coding in sequence.

[0077] The specific operations of various processing of S2 spatial coding are as follows:

[0078] Spherical coordinate transformation converts the longitude and latitude data (lat, lng) in the spatial coordinate information into spatial rectangular coordinates (x, y, z). The calculation formula is:

[0079]

[0080] Assume that the radius of the earth is r = 1, where θ represents latitude, φ represents longitude, and the value range of x, y, and z is in the interval [-1, 1].

[0081] Cube projection, construct a circumscribed cube with a side length of 1 outside the earth, realize the projection from the center of the sphere to the six faces of the cube, and obtain (f, u, v), where f is the identifier of the face, and (u, v) corresponds to the projection of the plane rectangular coordinates (x, y).

[0082] Projection correction, using quadratic transformation to adjust the projection coordinates (u, v), transforming (f, u, v) into (f, s, t), the calculation formula is:

[0083]

[0084] Where s, t are the corrected coordinates of u, v, and the range is [0, 1].

[0085] The projection correction uses a quadratic transformation to adjust the projection coordinates (u, v) to ensure that the size of each grid area is uniform to reduce the unevenness of space division.

[0086] Grid processing: divide the two-dimensional plane coordinates (u, v) into quadtree grid units according to the set level N and convert them into grid coordinates (i, j), where i and j represent the row and column numbers of the grid respectively, and the value range of i and j is [0,2 N -1].

[0087] Hilbert curve encoding processing uses the Hilbert curve to encode the grid coordinates and combines them with the surface identifier to form a CellID, which is finally converted into a decimal system to obtain a HID.

[0088] 3) Time coding:

[0089] The time information in the trajectory data processed in step 1) is converted into hexadecimal and combined to obtain TID.

[0090] The time represented in hexadecimal is an integer power of 2, which is more efficient when performing binary processing and improves computing efficiency.

[0091] 4) Index table design:

[0092] Design the Rowkey structure in the index table, combine the partition ID, HID, TID and vehicle identifier to generate the Rowkey, and ensure that each piece of data has a unique identifier;

[0093] The partition ID is the partition key, which divides and stores the trajectory data by day; FID is the vehicle ID.

[0094] 5) Build GeoMesa-HBase trajectory data storage model:

[0095] Combine the horizontal scalability of HBase with the GeoMesa spatiotemporal indexing tool to build an underlying data storage model.

[0096] 6) Build index table:

[0097] Based on the index table designed in step 4), GeoMesa is used to build the S2 index and store the data in the distributed database HBase. Figure 4 Schematic diagram of the principle for building an index table.

[0098] Before building the index, use the GeoMesa-HBase model API to traverse the main table data in the underlying storage;

[0099] Then build the index table. The value in the index table is the row key value in the main table.

[0100] The index table is built and stored in HBase, which is implemented through the GeoTools DataStore interface;

[0101] The mapping relationship between the index table and the main data table is constructed according to the designed Rowkey, and the data in the main table is located through the index table.

[0102] 7) Spatiotemporal query optimization strategy.

[0103] Figure 5 The following is a flow chart of spatiotemporal query. The query is performed through the spatiotemporal range query method. The spatiotemporal range query is a query scenario in which the time and space range conditions are pre-set to query data. First, the DataStore object is obtained according to the relevant parameters to realize the connection between GeoMesa, HBase and Zookeeper. The specific query steps are as follows:

[0104] Set query parameters according to the query conditions, determine the spatial range (Lat1, Lng1, Lat2, Lng2), time range (Stime, Etime), level N and table structure information, where Stime and Etime represent the query start time range and end time range, and level N corresponds to the spatial division in spatial coding;

[0105] Build a query object, use query parameters as query input, use ECQL syntax to filter and decompose query conditions, and build the final query object;

[0106] Determine the query scope and determine the query scope of the HBase database through the index of the query object;

[0107] Retrieve data and return query results.

[0108] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A storage and index construction method for trajectory data, characterized in that: The method comprises the following steps: 1) Preprocessing of original trajectory data: Get the GPS trajectory data of the taxi, which includes spatial coordinate information, time information and vehicle ID; Clean the trajectory data and retain the valid trajectory data; Unified data format; 2) S2 spatial coding: Perform S2 spatial encoding on the spatial coordinate information in the trajectory data processed in step 1) to generate a one-dimensional index value to obtain a HID; 3) Time coding: The time information in the trajectory data processed in step 1) is converted into hexadecimal and combined to obtain TID; 4) Index table design: Design the Rowkey structure in the index table, combine the partition ID, HID, TID and vehicle identifier to generate the Rowkey, and ensure that each piece of data has a unique identifier; The partition ID is the partition key, which divides and stores the trajectory data by day; FID is the vehicle ID; 5) Build GeoMesa-HBase trajectory data storage model: Combine the horizontal scalability of HBase with the GeoMesa spatiotemporal indexing tool to build an underlying data storage model; 6) Build index table: Based on the index table designed in step 4), use GeoMesa to build an S2 index and store the data in the distributed database HBase; 7) Spatiotemporal query optimization strategy: Set query parameters according to query conditions; Construct a query object; Determine the scope of the query; Retrieve data and return query results.

2. A method for storing and indexing trajectory data according to claim 1, characterized in that: In the step 1): The acquired trajectory data information is classified by day, and the acquisition frequency is 30s / time. The trajectory data information also includes whether there are passengers, speed and trajectory ID; Data cleaning includes removing missing values, removing redundant values, removing data outside the study area, removing values ​​with abnormal passenger instantaneous states, and processing trajectory drift data; Unifying the data format includes unifying the sorting of data and unifying the data format.

3. The method for storing and indexing trajectory data according to claim 1, characterized in that: In the step 2), the S2 space coding includes spherical coordinate transformation, cube projection, projection correction, gridding processing and Hilbert curve coding processing performed in sequence.

4. A method for storing and indexing trajectory data according to claim 3, characterized in that: The specific operations of various processing of the S2 space coding are as follows: Spherical coordinate transformation converts the longitude and latitude data (lat, lng) in the spatial coordinate information into spatial rectangular coordinates (x, y, z). The calculation formula is: Assume that the radius of the earth is r = 1, where θ represents latitude, φ represents longitude, and the range of x, y, and z is in the interval [-1, 1]; Cube projection, construct a circumscribed cube with a side length of 1 outside the earth, realize the projection from the center of the sphere to the six faces of the cube, and obtain (f, u, v), where f is the face identifier, and (u, v) corresponds to the projection of the plane rectangular coordinates (x, y); Projection correction, using quadratic transformation to adjust the projection coordinates (u, v), transforming (f, u, v) into (f, s, t), the calculation formula is: Where s, t are the corrected coordinates of u, v, and the range is [0, 1]; Grid processing: divide the two-dimensional plane coordinates (u, v) into quadtree grid units according to the set level N and convert them into grid coordinates (i, j), where i and j represent the row and column numbers of the grid respectively, and the value range of i and j is [0,2 N -1]; Hilbert curve encoding processing uses the Hilbert curve to encode the grid coordinates and combines them with the surface identifier to form a CellID, which is finally converted into a decimal system to obtain a HID.

5. The method for storing and indexing trajectory data according to claim 1, characterized in that: In the step 6), Before building the index, use the GeoMesa-HBase model API to traverse the main table data in the underlying storage; Then build the index table. The value in the index table is the row key value in the main table. The index table is built and stored in HBase, which is implemented through the GeoTools DataStore interface; The mapping relationship between the index table and the main data table is constructed according to the designed Rowkey, and the data in the main table is located through the index table.

6. The method for storing and indexing trajectory data according to claim 1, characterized in that: In the step 7), When setting query parameters according to query conditions, determine the spatial range (Lat1, Lng1, Lat2, Lng2), time range (Stime, Etime), level N and table structure information, where Stime and Etime represent the query start time range and end time range, and level N corresponds to the spatial division in spatial coding; When constructing a query object, use the query parameters as the query input and use the ECQL syntax to filter and decompose the query conditions; When determining the query scope, the query scope of the HBase database is determined by the index of the query object; Retrieve data and return query results.

Citation Information

Patent Citations

  • Spatial-temporal indexing method based on geographic grid and graph database

    CN117851695A

  • Method, system and device for retrieving spatio-temporal data and program product

    CN118839072A